Loading rankings…
Loading rankings…
Public operator evals vs public model evals
| Feature | SigRank | LMSYS Arena |
|---|---|---|
| What it ranks | AI operators (the humans using AI) | AI models (GPT, Claude, Gemini, etc.) |
| The question | Who is the best AI user? | Which AI model is best? |
| Method | Token telemetry — Yield (Υ) from session data | Human voting — pairwise preference (Elo) |
| Metric | Yield = (cache_read × output) / input² | Elo rating from human votes |
| Subjective? | No — measured from token counts | Yes — based on human preference |
| Data source | Operator's own session telemetry | Blind A/B votes on prompt outputs |
LMSYS Arena answers "which AI model is best?" — it ranks models by human preference in blind A/B tests. That's valuable for model selection.
SigRank answers "who is the best AI user?" — it ranks the humans who operate AI tools, by measuring how efficiently they use tokens. A great operator with a mediocre model can out-Yield a poor operator with the best model.
They're not competing — they're measuring different things. LMSYS measures the tool. SigRank measures the person wielding it.
LMSYS Arena ranks AI models by human preference in blind A/B votes (Elo). SigRank ranks the humans who use AI by Yield from token telemetry. LMSYS measures the tool; SigRank measures the person wielding it. A great AI user with a mediocre model can out-Yield a poor one with the best model.
No — they measure different subjects. LMSYS ranks models (GPT, Claude, Gemini); SigRank ranks the developers operating them. They're not competing — one evaluates the AI, the other evaluates the coder using the AI.
Yield (Υ) = (cache_read × output) / input² — token-cascade efficiency from real sessions. Works across Claude, GPT, Gemini, Cursor, Copilot, and any AI coding tool.
Visit signalaf.com/score to enroll and submit your token telemetry. SigRank will compute your Yield, your rank, and your operator class. Token counts only — never prompt content, never code.
The tool is the person.
Check my rank