Loading rankings…
Loading rankings…
The case for public AI operator evaluation
Models have Vals AI, LMSYS Arena, LiveBench. Operators have nothing. This creates an accountability gap where operator skill is invisible.
Public model evals created accountability for AI labs. Before public evals, model quality was marketing copy. After public evals, it was measurable. Operators deserve the same transparency \u2014 a public, comparable number that says "this is how well you use AI."
Yield (\u03a5) = (cache_read \u00d7 output) / input\u00b2. It measures cascade architecture \u2014 whether you're compounding signal or burning tokens. Not how much you spent, but what you got for what you spent.
The square on input means waste is non-linear. Operators who reuse context and extract real work score higher. Operators who burn fresh tokens on every turn score lower. Yield rewards the cascade, not the spend.
Companies hiring AI operators rely on take-home tests and vibes. Public operator evals replace vibes with data. A 6-month Yield history is more reliable than a 2-hour coding test.
A take-home test measures performance under artificial pressure for a single session. A Yield history measures performance across hundreds of real sessions \u2014 the actual cascade architecture the operator brings to real work. Hiring on vibes is how you end up with operators who talk about AI fluently but burn tokens inefficiently.
If Claude Code operators consistently out-Yield Cursor operators, that's public data. Tool selection becomes evidence-based.
Today, teams pick AI tools based on marketing, hype, and whichever influencer shouted loudest this week. Public operator evals turn tool selection into a measurable question \u2014 which tool produces the best cascade architectures in the hands of real operators doing real work.
Token counts only. Never prompt content. Never code. The evaluation is public; the work is private.
SigRank collects telemetry \u2014 input tokens, cache reads, output tokens. It never sees what you typed, what the model returned, or what code you shipped. The evaluation is public and comparable. The work stays yours.
Operator evals are public evaluations of AI users — the humans using AI. Unlike model evals that test AI models, operator evals measure how effectively a person uses AI, based on token telemetry from their real sessions. SigRank is the only platform running public operator evals.
Operator evals create accountability for AI users. Before public evals, developer skill with AI was vibes. After public evals, it's a measurable, comparable, public number. The same transparency that public model evals brought to AI labs, public operator evals bring to AI coders.
Models have public evals (Vals AI, LMSYS Arena, LiveBench). AI users don't. This creates an accountability gap where developer skill with AI is invisible. Public operator evals close that gap by benchmarking the human, not the model.
Companies hiring AI developers can look at public Yield scores instead of take-home assignments. A 6-month Yield history is more reliable than a 2-hour coding test. The leaderboard is the portfolio. Works across Claude, GPT, Cursor, Copilot, and any AI coding tool.
SigRank collects token counts only — never prompt content, never code, never conversation. The evaluation is public; the work is private. Your Yield score is on the leaderboard; your code is not.
Yield (Υ) = (cache_read × output) / input². It measures token-cascade efficiency — whether signal is compounding or tokens are being burned. Works across any AI platform — Claude, GPT, Gemini, Cursor, Copilot, or any coding agent.
Get your public operator eval.
Measure your Yield. See your rank. Join the public evaluation.
Check my rank