Loading rankings…
Loading rankings…
Public LLM operator evals vs public LLM model evals
| Feature | SigRank | Vals AI |
|---|---|---|
| What it evaluates | AI operators (the humans using AI) | AI models (GPT, Claude, Gemini, etc.) |
| The question | Who is the best AI operator? | Which AI model is best? |
| Eval type | Public operator evals — telemetry-based | Public model evals — benchmark-based |
| Method | Token telemetry \u2014 Yield (\u03a5) from session data | Standardized benchmark suites (MMLU, HumanEval, etc.) |
| Metric | Yield = (cache_read \u00d7 output) / input\u00b2 | Benchmark scores (accuracy, pass rate, etc.) |
| Subjective? | No \u2014 measured from token counts | No \u2014 measured from benchmark results |
| Data source | Operator's own session telemetry (privacy-preserving) | Standardized test prompts run against models |
| Public? | Yes \u2014 public leaderboard, public methodology | Yes \u2014 public results, public benchmarks |
Vals AI runs public model evals — standardized benchmark suites that measure how well an AI model performs on reasoning, coding, math, and knowledge tasks. The subject is the model. The question is "which model is best?"
SigRank runs public operator evals —telemetry-based evaluations that measure how effectively a human operator uses AI. The subject is the person. The question is "who is the best AI operator?"
They're complementary, not competing. Vals AI tells you which model to use. SigRank tells you how well you're using it. A great operator with a mediocre model can out-Yield a poor operator with the best model \u2014 because Yield measures the cascade architecture, not the model's raw capability.
Public model evals (Vals AI, LMSYS Arena, LiveBench) created accountability for AI labs. Before public evals, model quality was marketing copy. After public evals, it was measurable.
Public operator evals do the same for AI users. Before SigRank, operator skill was vibes \u2014 "they seem productive." After SigRank, it's Yield \u2014 a measurable, comparable, public number. The same transparency that public model evals brought to AI labs, public operator evals bring to AI operators.
Vals AI runs public model evals — standardized benchmarks that measure how well an AI model performs. SigRank runs public operator evals — telemetry-based evaluations that measure how effectively a human uses AI. Vals AI evaluates the model (GPT, Claude, Gemini); SigRank evaluates the person wielding it.
No, they're complementary. Vals AI tells you which model to use. SigRank tells you how well you're using it. A great AI user with a mediocre model can out-Yield a poor AI user with the best model — because Yield measures the cascade architecture, not the model's raw capability.
SigRank evaluates the human — the developer, coder, or AI user — and their token-cascade efficiency (Yield). Vals AI evaluates the AI model — its benchmark performance. Different subjects, different metrics, different questions. SigRank is the only platform running public evals for AI users.
SigRank uses Yield (Υ) = (cache_read × output) / input² — token-cascade efficiency from real sessions. Vals AI uses benchmark scores (accuracy, pass rate) from standardized test prompts. SigRank measures the person; Vals AI measures the model.
Yes. Both Vals AI and SigRank publish public results with public methodology. Vals AI publishes model benchmark scores publicly. SigRank publishes AI user Yield scores on a public leaderboard.
Vals AI ranks AI models — Claude, GPT, Gemini, and others — by benchmark performance. SigRank ranks the humans using those models — developers, coders, AI users — by Yield. Vals AI tells you which model is best; SigRank tells you who is the best at using it.
Vals AI evaluates the model. SigRank evaluates the operator.
Public evals for the humans wielding AI.
Get your operator eval