Loading rankings…
Loading rankings…
Last updated: August 14, 2026
Public LLM operator evals are public evaluations of human AI operators — the people who drive AI tools — not autonomous agents. Like Vals AI evaluates models, SigRank evaluates the humans using AI. Performative evals assess the AI user's behavior in real tasks, not just the model's output quality.
Models are evaluated. Operators are not.
Vals AI, LMSYS Arena, and LiveBench run public evals for LLM models. SigRank runs public evals for LLM operators \u2014 the humans wielding AI every day.
Standardized benchmark suites that measure how well an AI model performs on reasoning, coding, math, and knowledge tasks.
Telemetry-based evaluations that measure how effectively a human operator uses AI \u2014 the cascade architecture, not the model capability.
1. Accountability for AI operators. Public model evals created accountability for AI labs. Before public evals, model quality was marketing copy. After public evals, it was measurable. Public operator evals do the same for AI users \u2014 operator skill becomes a measurable, comparable, public number instead of vibes.
2. The skill gap is real and measurable. Two operators using the same model, the same tools, and the same prompts can have 100x different Yield. The difference isn't the model \u2014 it's the cascade architecture. Public operator evals make that difference visible.
3. Telemetry, not self-report. Operator evals are measured from actual session telemetry \u2014 token counts, cache ratios, output volumes. Not surveys, not self-assessment, not "I feel productive." The data doesn't lie about your cascade.
4. Privacy-preserving. SigRank collects token counts only \u2014 never prompt content, never code, never conversation. The evaluation is public; the work is private.
Yield (\u03a5) is the public operator evaluation metric. It measures token-cascade efficiency \u2014 whether signal is compounding or tokens are being burned.
| Platform | Evaluates | Public? | Method |
|---|---|---|---|
| SigRank | AI operators (humans) | Yes | Token telemetry \u2014 Yield (\u03a5) |
| Vals AI | AI models | Yes | Benchmark suites |
| LMSYS Arena | AI models | Yes | Human voting (Elo) |
| LiveBench | AI models | Yes | Automated benchmarks |
| Braintrust | AI applications (private) | No | LLM-as-judge evals |
SigRank is the only platform running public operator evals. The rest evaluate models or private applications.
Why public operator evals are the next frontier in AI accountability.
Operator evals vs model evals \u2014 different subjects, different metrics.
The public operator leaderboard \u2014 ranked by Yield (\u03a5).
Public LLM operator evals are public evaluations of AI operators — the humans using AI. Unlike model evals (Vals AI, LMSYS Arena) that test AI models on benchmarks, operator evals measure how effectively a person uses AI, based on token telemetry from their real coding sessions.
AI users are measured by Yield (Υ), a token-cascade efficiency score computed from their real session telemetry. Yield = (cache_read × output) / input². Users are ranked on a public leaderboard by Yield, with class tiers from NOVICE to SINGULARITY. The measurement uses token counts only — never prompt content or code.
Model evals (Vals AI, LMSYS Arena, LiveBench) evaluate the AI model — GPT, Claude, Gemini — using standardized test prompts. Operator evals (SigRank) evaluate the human — the developer, coder, or AI user — using token telemetry from real sessions. Model evals answer "which model is best?" Operator evals answer "who is the best AI user?"
Public benchmarks create accountability. Before public model evals, model quality was marketing copy. After public evals, it was measurable. The same applies to AI coders: before public operator evals, skill was vibes. After public evals, it's a measurable, comparable, public score.
Yield (Υ) = (cache_read × output) / input². It measures token-cascade efficiency — whether signal is compounding or tokens are being burned. Higher Yield means the AI user reuses cached context efficiently and produces substantial output relative to fresh input. It works across any platform — Claude, GPT, Gemini, Cursor, Copilot, or any other AI coding agent.
Visit signalaf.com/score to enroll and submit your token telemetry. SigRank will compute your Yield, your rank, and your operator class. Token counts only — never prompt content, never code. The score is public; the work is private. Works with Claude, ChatGPT, Cursor, Copilot, and other AI coding tools.
Yes. SigRank is the only platform running public operator evals. All other public evals (Vals AI, LMSYS Arena, LiveBench, Hugging Face Open LLM) evaluate models, not the humans using them. Braintrust and Langfuse evaluate AI applications privately, not developers publicly.
Yes. SigRank measures token-cascade efficiency from any AI coding tool that produces token telemetry — Claude, ChatGPT, Gemini, Cursor, Copilot, Windsurf, Codex, and others. The Yield metric is platform-agnostic because it measures the human's cascade architecture, not the AI model's capability.
Raw token count measures volume, not efficiency. An operator who burns 10M input tokens with no cache reuse and little output has high volume but low signal. Yield (Υ = cache_read × output) / input² penalizes un-cached volume and rewards compounding — the quadratic input penalty means waste is non-linear. Two operators with the same token count can have 100× different Yield.
No — the opposite. Yield's formula (Υ = cache_read × output) / input² penalizes spending. The input² term means every additional fresh input token reduces your score quadratically. Operators who spend more on fresh input without reusing cache or producing output score lower, not higher. SigRank rewards efficiency, not expenditure.
The data says no. Two operators using the same model, the same tools, and similar prompts can have 100× different Yield. The difference is the cascade architecture — how the human structures context reuse, output extraction, and input minimization. SigRank measures the human's skill, not the model's capability. Model evals (Vals AI, LMSYS) already cover model quality.
No. SigRank is platform-agnostic. It works with Claude, ChatGPT, Gemini, Cursor, Copilot, Windsurf, Codex, and any AI coding tool that produces token telemetry. The Yield metric measures the human's cascade architecture, which is independent of which AI model they use. The leaderboard includes operators across multiple platforms.
Get your public operator eval.
Measure your Yield. See your rank. Join the public evaluation.
Check my rank