TonkaToyXL

BENCHMARK DECODER // PROOF BEFORE RANK

Compare three AI models without hiding the tradeoffs.

A benchmark is a measurement, not a verdict. Start with intelligence, cost, and speed, then inspect latency, response time, context, and the conditions behind every number.

Looking for GPT-6 Astra? This comparator preserves the August 29 snapshot. Open the sourced Astra field guide and test prompts →

INDEPENDENT SIGNAL // THREE-WAY MODEL CHECK

Put three models on the same workbench.

Select any three current model variants. Reasoning levels stay separate, and the table marks the strongest measured value in each row without flattening six tradeoffs into one universal rank.

IntelligenceClaude Opus 5 (max)

Cost per taskGemini 3.7 Flash (high)

Output speedGemini 3.7 Flash (high)

Artificial Analysis benchmark comparison for the three selected models
Measured factorOpenAIGPT-5.6 Sol (max)AnthropicClaude Opus 5 (max)GoogleGemini 3.7 Flash (high)
IntelligenceAA Intelligence Index · higher is better6163Leads selected56
Cost per taskWeighted AA Intelligence Index task · lower is better$0.95$2.34$0.40Leads selected
Output speedMedian output tokens per second · higher is better75 t/s55 t/s324 t/sLeads selected
First-chunk latencySeconds to first chunk · lower is better153.40s60.19s9.52sLeads selected
Total responseSeconds for the measured response · lower is better160.03s69.22s11.06sLeads selected
Context windowMaximum combined token window · higher is better1M tokensLeads selected1M tokensLeads selected1M tokensLeads selected

Snapshot checked August 29, 2026. Reasoning levels use the source's exact labels; unlisted levels are not inferred. Values change as models, providers, and benchmark methods change. Cost per task is Artificial Analysis' weighted Intelligence Index task cost, not a subscription price.

CABIN TEST LAB // EVIDENCE LANES

The selector above displays a dated independent Artificial Analysis snapshot. The TonkaToyXL Test Lab remains separate: our latest videos report bounded cabin tests with their own prompts, refinements, and limits. Those results do not change this independent snapshot.

See the latest cabin tests →

FIVE-PART BENCHMARK READHIGHER IS NOT ALWAYS BETTER

01 // Exact model

A family name is not enough. Confirm the version, reasoning mode, tool setup, and any special configuration.

REQUIRED CONTEXT

02 // Exact benchmark

Versions change. Scores from different harnesses or test revisions may not be comparable.

REQUIRED CONTEXT

03 // Who ran it

Vendor-reported results and independent evaluations answer different evidence questions.

REQUIRED CONTEXT

04 // Direction and scale

Check whether higher or lower is better and whether the chart begins at zero.

REQUIRED CONTEXT

05 // Conditions and limits

Budget, retries, time limits, judge models, contamination, and tool access can change the outcome.

REQUIRED CONTEXT

Do not blend evidence until it looks like certainty.

Vendor-reported

Useful for understanding what the maker claims and how it configured the test. It is not independent validation.

Independent test

Useful when the evaluator publishes the model version, harness, conditions, costs, and limitations. Independence does not automatically make a method good.

Your task test

The final decision layer. Use work you understand, record the setup, compare outputs blind when practical, and count the time needed to correct each result.

What a score cannot tell you alone

Whether the model follows your style, protects your data, works with your tools, stays inside budget, or produces something you can trust without excessive review.