Compare three AI models without hiding the tradeoffs.
A benchmark is a measurement, not a verdict. Start with intelligence, cost, and speed, then inspect latency, response time, context, and the conditions behind every number.
Select any three current model variants. Reasoning levels stay separate, and the table marks the strongest measured value in each row without flattening six tradeoffs into one universal rank.
IntelligenceClaude Opus 5 (max)
Cost per taskGemini 3.7 Flash (high)
Output speedGemini 3.7 Flash (high)
Artificial Analysis benchmark comparison for the three selected models
Measured factor
OpenAIGPT-5.6 Sol (max)
AnthropicClaude Opus 5 (max)
GoogleGemini 3.7 Flash (high)
IntelligenceAA Intelligence Index · higher is better
61
63Leads selected
56
Cost per taskWeighted AA Intelligence Index task · lower is better
$0.95
$2.34
$0.40Leads selected
Output speedMedian output tokens per second · higher is better
75 t/s
55 t/s
324 t/sLeads selected
First-chunk latencySeconds to first chunk · lower is better
153.40s
60.19s
9.52sLeads selected
Total responseSeconds for the measured response · lower is better
160.03s
69.22s
11.06sLeads selected
Context windowMaximum combined token window · higher is better
1M tokensLeads selected
1M tokensLeads selected
1M tokensLeads selected
Snapshot checked August 29, 2026. Reasoning levels use the source's exact labels; unlisted levels are not inferred. Values change as models, providers, and benchmark methods change. Cost per task is Artificial Analysis' weighted Intelligence Index task cost, not a subscription price.
CABIN TEST LAB // EVIDENCE LANES
The selector above displays a dated independent Artificial Analysis snapshot. The TonkaToyXL Test Lab remains separate: our latest videos report bounded cabin tests with their own prompts, refinements, and limits. Those results do not change this independent snapshot.
FIVE-PART BENCHMARK READHIGHER IS NOT ALWAYS BETTER
01 // Exact model
A family name is not enough. Confirm the version, reasoning mode, tool setup, and any special configuration.
REQUIRED CONTEXT
02 // Exact benchmark
Versions change. Scores from different harnesses or test revisions may not be comparable.
REQUIRED CONTEXT
03 // Who ran it
Vendor-reported results and independent evaluations answer different evidence questions.
REQUIRED CONTEXT
04 // Direction and scale
Check whether higher or lower is better and whether the chart begins at zero.
REQUIRED CONTEXT
05 // Conditions and limits
Budget, retries, time limits, judge models, contamination, and tool access can change the outcome.
REQUIRED CONTEXT
Do not blend evidence until it looks like certainty.
Vendor-reported
Useful for understanding what the maker claims and how it configured the test. It is not independent validation.
Independent test
Useful when the evaluator publishes the model version, harness, conditions, costs, and limitations. Independence does not automatically make a method good.
Your task test
The final decision layer. Use work you understand, record the setup, compare outputs blind when practical, and count the time needed to correct each result.
What a score cannot tell you alone
Whether the model follows your style, protects your data, works with your tools, stays inside budget, or produces something you can trust without excessive review.