Loading rankings…
Loading rankings…
Different subjects, different metrics, different questions
What Vals AI, LMSYS Arena, LiveBench do. Standardized benchmarks. Subject = the model. Question = "which model is best?" Method = test prompts. Output = accuracy/Elo scores.
Model evals run standardized test prompts against an AI model and measure how well it performs. The subject is the model. The question is "which model is best?" The method is a benchmark suite. The output is a score \u2014 accuracy, pass rate, Elo rating. Model evals created accountability for AI labs. Before them, model quality was marketing copy.
What SigRank does. Telemetry from real sessions. Subject = the human. Question = "who is the best AI operator?" Method = token cascade analysis. Output = Yield (\u03a5).
Operator evals measure telemetry from real work sessions \u2014 not test prompts, not benchmarks. The subject is the human operator. The question is "who is the best AI operator?" The method is token cascade analysis. The output is Yield (\u03a5), a single comparable number that reflects how well the operator compounds signal across a session.
| Feature | Model Evals | Operator Evals |
|---|---|---|
| Subject | The AI model | The human operator |
| Question | Which model is best? | Who is the best AI operator? |
| Method | Standardized test prompts | Token cascade analysis |
| Metric | Accuracy / Elo scores | Yield (\u03a5) |
| Data source | Benchmark suites run against models | Telemetry from real work sessions |
| Public? | Yes \u2014 public results, public benchmarks | Yes \u2014 public leaderboard, public methodology |
| Examples | Vals AI, LMSYS Arena, LiveBench | SigRank |
Model evals tell you which model to use. Operator evals tell you how well you're using it. A great operator with a mediocre model can out-Yield a poor operator with the best model.
They answer different questions and they're both important. Model evals are about the tool. Operator evals are about the craft. You need both to understand AI-assisted work \u2014 a great tool in unskilled hands produces mediocre results, and a mediocre tool in skilled hands can still compound signal efficiently.
Model evals became infrastructure. Operator evals will too. The gap: nobody was evaluating the human. SigRank fills that gap.
Vals AI, LMSYS Arena, and LiveBench built the public evaluation layer for models. That layer is now infrastructure \u2014 every AI lab checks it, every buyer references it. The same thing will happen for operator evals. The gap was simple: nobody was evaluating the human wielding the model. SigRank is the public evaluation layer for operators.
Model evals evaluate AI models (GPT, Claude, Gemini) using standardized benchmarks. Operator evals evaluate the humans using AI using token telemetry from real sessions. The subject is different: one evaluates the tool, the other evaluates the person wielding it.
No, they're complementary. Model evals tell you which model to use. Operator evals tell you how well you're using it. A great AI user with a mediocre model can out-Yield a poor AI user with the best model.
Vals AI, LMSYS Arena, LiveBench, and Hugging Face Open LLM Leaderboard are all public model evals. They run standardized test prompts against AI models and publish the results.
SigRank is the only platform running public operator evals. All other public evals (Vals AI, LMSYS Arena, LiveBench) evaluate models, not the humans using them.
Models are evaluated publicly. Applications are evaluated privately (Braintrust, Langfuse). But AI users — the developers, coders, and people using AI — had no public evaluation layer until SigRank. That's the missing layer.
SigRank ranks developers — the humans using AI. Vals AI and LMSYS Arena rank the models (Claude, GPT, Gemini). SigRank is the missing layer: public evals for the person, not the platform.