Model tracker
The leading models compared on independent benchmarks — intelligence, evals, price and speed — refreshed every morning, with day-on-day movement and the longer trend.
Editor's notes: the top five in practice
Based on the 12 June 2026 capture. Benchmarks measure benchmarks — always trial a model on your own work before committing.
Claude Fable 5 (Adaptive Reasoning, Max Effort)
Best for: the hardest problems, when quality outranks everything.
The leader on nearly every measure we track — intelligence (64.9), agentic tool use (tau², 0.99), autonomous terminal work, and the punishing Humanity's Last Exam — but it makes you pay twice for it: roughly double the price of its rivals ($10 in / $50 out per million tokens) and over a minute before the first word arrives at this effort setting. A deep-work specialist, not a conversation partner. Worth noting: independent testers have rated its everyday coding more modestly than its benchmark scores suggest — see our 12 June transmission.
Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
Best for: production agents.
The balance point of the field: second-best agentic scores (tau² 0.94), half Fable's price ($5 / $25), and a 19-second start instead of a minute. If you are building systems that browse, call tools and work unattended, this is where capability per pound currently peaks.
GPT-5.5 (xhigh)
Best for: research, analysis and long documents.
The top scorer on graduate-level science questions (GPQA 0.935) and on long-context reasoning — the measure that matters when you feed a model a 200-page report. Strong second on coding. Slow to start (~45 seconds), so suited to considered work rather than quick exchanges.
GPT-5.5 (high)
Best for: interactive development.
Nearly all of xhigh's ability at the same price, with the wait cut to about 14 seconds — the difference between a coding assistant you talk to and one you queue for. The sensible default in the GPT-5.5 family unless you need the last benchmark point.
Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
Best for: proven, role-specific deployments.
A step behind its successor on every measure, at the same price — but the fastest first response of the five (under 11 seconds) and months of production track record. The conservative choice for customer-facing assistants and established workflows that prize predictability over peak scores.
Data: Artificial Analysis, captured 12 June 2026. Prices are per million tokens. We have no affiliation with any model maker; nothing here is a recommendation to buy.
No benchmark snapshot captured yet — the first lands with tomorrow morning's pipeline run.
Data: Artificial Analysis — captured 2026-06-20
Capabilities
Which frontier models support which capabilities. A tick means the laboratory ships the feature in its main product or API. Showing preview data until the live pipeline is connected.
| Model | Long-running agents | Code execution | Computer use | Voice | Open weights | Notes |
|---|---|---|---|---|---|---|
Claude Fable 5 Anthropic · 2026 | ● | ● | ● | — | — | Frontier tier; built for extended autonomous tasks. |
GPT-5.5 OpenAI · 2026 | ● | ● | ● | ● | — | Includes agentic coding and operator-style computer control. |
Gemini 3.1 Pro Google DeepMind · 2026 | ● | ● | ● | ● | — | Deep integration with Google Workspace and search. |
DeepSeek V3.2 DeepSeek · 2025 | — | ● | — | — | ● | Leading open-weights option at frontier-adjacent quality. |
Grok 4.1 xAI · 2025 | ● | ● | — | ● | — | Strong human-preference scores on community leaderboards. |
Mistral Large 3 Mistral · 2025 | — | ● | — | — | ● | European laboratory; strong open-weights line. |