Frontier Model Comparison
A point-in-time benchmark snapshot for comparing leading AI models.
Data reviewed: August 29, 2026
Select models to compare
| Spec | GPT-5.6 Sol | Claude Opus 5 | Gemini 3.7 Flash | Grok 4.6 |
|---|---|---|---|---|
| Released | Jul 9, 2026 | Jul 24, 2026 | Aug 13, 2026 | Aug 12, 2026 |
| API model ID | gpt-5.6-sol | claude-opus-5 | gemini-3.7-flash | grok-4.6 |
| Context | 1.05M | 1M | 1M | 500K |
| Max output | 128K | 128K | 64K | No fixed text limit |
| Knowledge / training cutoff | February 2026 | May 2026 | March 2026 | February 2026 |
| Published benchmark results | 12 of 32 | 9 of 32 | 10 of 32 | 4 of 32 |
| Input / 1M tokens | $5.00 | $5.00 | $0.75 | $2.00 |
| Output / 1M tokens | $30.00 | $25.00 | $3.75 | $6.00 |
GPT-5.6 Sol: Prompts above 272K tokens use OpenAI's long-context rate.
Gemini 3.7 Flash: Introductory price through December 31, 2026; standard price is $1.50/$7.50.
Grok 4.6: Base rate below 200K input tokens; higher long-context rates apply.
| Benchmark | GPT-5.6 Sol | Claude Opus 5 | Gemini 3.7 Flash | Grok 4.6 |
|---|---|---|---|---|
GPQA Diamond PhD-level science reasoning | 94.6% | — | — | — |
GDPval Knowledge work across 44 occupations | — | — | — | — |
MMMU Pro Multimodal understanding | 83.0% | — | — | — |
ARC-AGI-2 Novel problem-solving | — | 90.4% | — | — |
HLE High-level exam performance | — | 64.7% | 53.6% | — |
Agents' Last Exam Long-running professional workflows across 55 fields | 52.7% | — | 26.3% | — |
GDPval-AA v2 Agentic knowledge-work performance (Elo) | 1,748 | 1,861 | 1,525 | 1,753 |
AA Intelligence Index Artificial Analysis composite intelligence index | 58.9 | — | 56.0 | 61.0 |
SWE-bench Verified Real-world software engineering | — | 96.0% | — | — |
SWE-bench Pro Advanced software engineering | 64.6% | 79.2% | — | — |
Terminal-Bench 2.0 Terminal-based coding tasks | — | — | — | — |
Terminal-Bench 2.1 Updated terminal-based coding tasks | 88.8% | — | 85.8% | — |
Terminal-Bench 3.0 Third-generation terminal agent benchmark | — | — | 14.9% | 26.0% |
DeepSWE v1.1 Long-horizon engineering in real repositories | 72.7% | 68.8% | 65.3% | 65.9% |
MLE Bench Lite Autonomous ML research competitions | — | — | — | — |
VIBE-Pro Full project delivery (web, mobile, simulation) | — | — | — | — |
OSWorld Verified Computer use tasks | — | — | — | — |
OSWorld 2.0 Updated autonomous computer-use tasks | 62.6% | 70.6% | 47.9% | — |
WebArena Verified Web-based agent tasks | — | — | — | — |
BrowseComp Web browsing comprehension | 90.4% | 90.8% | — | — |
Toolathlon Tool use proficiency | 58.0% | — | — | — |
Toolathlon Verified Verified tool-use proficiency tasks | — | — | — | — |
AutomationBench End-to-end business workflow automation | 18.1% | 26.0% | 30.4% | — |
CyberGym Source-grounded vulnerability discovery | — | — | — | — |
MCP Atlas MCP tool integration | — | — | — | — |
PinchBench OpenClaw agent brain evaluation (Kilo.ai) | — | — | — | — |
ClawEval Real-world agent tasks across work and life | — | — | — | — |
Tau-bench Tool-use reliability under real conditions | — | — | — | — |
LMArena Overall Crowdsourced human preference (Elo) | — | — | — | — |
LMArena Coding Coding preference (Elo) | — | — | — | — |
WebDev Arena Human preference for generated web applications (Elo) | — | — | 1,588 | — |
Non-Hallucination Rate AA-Omniscience: % of non-correct responses that are abstentions (not wrong answers) | — | — | — | — |
Benchmark values come from provider model cards, release disclosures, and named independent leaderboards. Harnesses, reasoning effort, and test versions can vary. A dash means we found no comparable published result, not a score of zero. Versioned benchmarks stay in separate rows.
Sources
OpenAI GPT-5.6 ↗OpenAI GPT-5.6 API models ↗OpenAI GPT-5.5 ↗Anthropic current models ↗Anthropic Claude Fable 5 ↗Anthropic Claude Opus 5 system card ↗Anthropic Claude Sonnet 5 system card ↗Google Gemini 3.7 Flash model card ↗Google Gemini 3.5 Flash-Lite model card ↗xAI Grok 4.6 ↗Z.ai GLM-5.3 ↗Kimi K3 ↗DeepSeek V4 ↗DeepSeek API pricing ↗Alibaba Qwen3.8 Max ↗MiniMax M3 ↗OpenAI GPT-5.4 ↗Anthropic Claude Opus 4.6 ↗Google Gemini 3 Flash ↗Google DeepMind Gemini 3.1 Pro ↗Evolink AI ↗Digital Applied ↗Arena.ai Leaderboard ↗vals.ai SWE-bench ↗Artificial Analysis AA-Omniscience ↗Suprmind Hallucination Reference ↗Onyx Self-Hosted LLM Leaderboard ↗PinchBench ↗Google Gemma 4 ↗Google DeepMind Gemma 4 ↗