How to read these scores
The Epoch Capabilities Index combines many benchmarks into one general-capability scale. The numbers in brackets are its confidence interval: when two models’ intervals overlap, treat them as roughly equal rather than declaring a winner.
Each benchmark score is the best result Epoch AI recorded for that model, usually at its highest reasoning setting. Higher reasoning settings produce more output tokens, so a model can cost more in practice than its per-token price suggests.
The benchmarks
- GPQA Diamond (science reasoning): Graduate-level biology, chemistry and physics questions designed to be hard to look up.
- FrontierMath (Tiers 1–3) (advanced math): Original, unpublished math problems ranging from advanced undergraduate to research level.
- SimpleQA Verified (factual accuracy): Short fact-seeking questions that reward correct answers over confident guesses.
- ARC-AGI-2 (abstract reasoning): Visual pattern puzzles that are easy for people but hard for AI systems.
- DeepSWE (coding): Agentic software engineering tasks that require changing real codebases.
- APEX-Agents (agentic work): Long, multi-step professional tasks completed by an AI agent.
A dash means Epoch AI has not published a score for that model on that benchmark. Benchmarks measure narrow skills; test candidates on your own tasks before committing. To turn these numbers into a monthly bill, use the cost calculator or compare two models.
Data: Epoch AI, “Capabilities & Benchmarking”, epoch.ai/benchmarks, licensed CC BY 4.0. Prices from each provider’s official pricing page.