Evidence before ranking.
The rules used to make model comparisons useful without pretending they are universal truth.
Benchmark-specific ranks
A model ranks only within a named evaluation; unrelated tests are never averaged.
Independent versus reported
Independent measurements and provider claims must carry explicit labels.
Missing means missing
Unpublished results remain blank—we never interpolate or estimate scores.
Versions remain attached
Benchmark version, reasoning effort, harness, API and date remain attached.
No false precision
Small score differences may not be practically meaningful; test your own workload.
Correction policy
Source corrections trigger a reviewed snapshot update and new evidence date.