EVIDENCE, NOT HYPEMeasure the model.
Measure the model.
Keep the context.
Independent benchmark signals organized without inventing one universal winner. Every score keeps its evaluation, version, date, and source attached.
9 MODELS4 MEASURES0 ESTIMATES
WORKSPACE / PUBLISHED SNAPSHOTChange the measure.
Change the measure.
Change the answer.
Updated August 20, 2026
ACTIVE MEASURE
Intelligence Index
A synthesis of agentic, coding, scientific-reasoning, and general evaluations. Methodology ↗
RANKMODELACCESSSCORERELATIVE
01Claude Opus 5AnthropicProprietary61index↗02Claude Fable 5AnthropicProprietary60index↗03GPT-5.6 SolOpenAIProprietary59index↗04Kimi K3Moonshot AIOpen weights57index↗05GPT-5.5OpenAIProprietary55index↗06Muse Spark 1.2MetaProprietary54index↗07Grok 4.5SpaceXAIProprietary54index↗08Kimi K2.6Moonshot AIOpen weights54index↗09North Mini CodeCohereOpen weights27.6index↗Ranks apply only to the selected evaluation and cited snapshot. Missing scores are not estimated. Harness, reasoning effort, provider, and version can materially change results.
METHOD / 01
No universal score.
Unrelated tests are never averaged into an invented ranking.
EVIDENCE / 02
Source attached.
Every number links to the underlying independent evaluation.
LIMITS / 03
READ THE FULL METHODOLOGY ↗Conditions matter.
Harnesses, effort settings, APIs, and dates affect outcomes.