Frontier Model Comparison

A point-in-time benchmark snapshot for comparing leading AI models.

Data reviewed: August 29, 2026

Select models to compare

SpecGPT-5.6 SolClaude Opus 5Gemini 3.7 FlashGrok 4.6
ReleasedJul 9, 2026Jul 24, 2026Aug 13, 2026Aug 12, 2026
API model IDgpt-5.6-solclaude-opus-5gemini-3.7-flashgrok-4.6
Context1.05M1M1M500K
Max output128K128K64KNo fixed text limit
Knowledge / training cutoffFebruary 2026May 2026March 2026February 2026
Published benchmark results12 of 329 of 3210 of 324 of 32
Input / 1M tokens$5.00$5.00$0.75$2.00
Output / 1M tokens$30.00$25.00$3.75$6.00

GPT-5.6 Sol: Prompts above 272K tokens use OpenAI's long-context rate.

Gemini 3.7 Flash: Introductory price through December 31, 2026; standard price is $1.50/$7.50.

Grok 4.6: Base rate below 200K input tokens; higher long-context rates apply.

BenchmarkGPT-5.6 SolClaude Opus 5Gemini 3.7 FlashGrok 4.6
GPQA Diamond
PhD-level science reasoning
94.6%
GDPval
Knowledge work across 44 occupations
MMMU Pro
Multimodal understanding
83.0%
ARC-AGI-2
Novel problem-solving
90.4%
HLE
High-level exam performance
64.7%53.6%
Agents' Last Exam
Long-running professional workflows across 55 fields
52.7%26.3%
GDPval-AA v2
Agentic knowledge-work performance (Elo)
1,7481,8611,5251,753
AA Intelligence Index
Artificial Analysis composite intelligence index
58.956.061.0
SWE-bench Verified
Real-world software engineering
96.0%
SWE-bench Pro
Advanced software engineering
64.6%79.2%
Terminal-Bench 2.0
Terminal-based coding tasks
Terminal-Bench 2.1
Updated terminal-based coding tasks
88.8%85.8%
Terminal-Bench 3.0
Third-generation terminal agent benchmark
14.9%26.0%
DeepSWE v1.1
Long-horizon engineering in real repositories
72.7%68.8%65.3%65.9%
MLE Bench Lite
Autonomous ML research competitions
VIBE-Pro
Full project delivery (web, mobile, simulation)
OSWorld Verified
Computer use tasks
OSWorld 2.0
Updated autonomous computer-use tasks
62.6%70.6%47.9%
WebArena Verified
Web-based agent tasks
BrowseComp
Web browsing comprehension
90.4%90.8%
Toolathlon
Tool use proficiency
58.0%
Toolathlon Verified
Verified tool-use proficiency tasks
AutomationBench
End-to-end business workflow automation
18.1%26.0%30.4%
CyberGym
Source-grounded vulnerability discovery
MCP Atlas
MCP tool integration
PinchBench
OpenClaw agent brain evaluation (Kilo.ai)
ClawEval
Real-world agent tasks across work and life
Tau-bench
Tool-use reliability under real conditions
LMArena Overall
Crowdsourced human preference (Elo)
LMArena Coding
Coding preference (Elo)
WebDev Arena
Human preference for generated web applications (Elo)
1,588
Non-Hallucination Rate
AA-Omniscience: % of non-correct responses that are abstentions (not wrong answers)

Benchmark values come from provider model cards, release disclosures, and named independent leaderboards. Harnesses, reasoning effort, and test versions can vary. A dash means we found no comparable published result, not a score of zero. Versioned benchmarks stay in separate rows.