TonkaToyXL

CABIN WATCH DESK // TONKATOYXL

Watch AI get practical.

Long-form explainers, short experiments, model reality checks, and source-backed technology stories, packaged for people who want the evidence.

LATEST VERIFIED RELEASES

Find your next rabbit hole.

Subscribe on YouTube

Updated September 5, 2026 from the official TonkaToyXL channel. Publication titles are preserved; dates are shown in UTC. Release verification confirms the upload, not every claim inside it.

41 stories in the cabin library

Short

MODEL COMPARISON

PUBLIC RELEASE VERIFIED All levels

Four AI Models Build the Cabin in Blender | TonkaToyXL

The Blender challenge in miniature: compare the final cabin renders, with unavailable results kept visible.

Watch on YouTube
Short

MODEL COMPARISON

PUBLIC RELEASE VERIFIED All levels

Four AI Models Build a Skybridge Cabin | TonkaToyXL

A quick look at the explorable skybridge challenge and why a failed final build stays a failed result.

Watch on YouTube
Short

MODEL COMPARISON

PUBLIC RELEASE VERIFIED All levels

Four AI Models Build a Cabin Kart Game | TonkaToyXL

A snowy kart-game challenge with one first attempt and two refinements for each model.

Watch on YouTube
Video

MODEL COMPARISON

PUBLIC RELEASE VERIFIED All levels

Is Astra Worth It? Four AI Models Build the Test Cabin

Astra, Sol, Qwen3.8 Max, and DeepSeek V4 Pro tackle three cabin builds. Inspect the outputs, refinements, regressions, and cost limits.

Watch on YouTube
Short

ASTRA TEST

PUBLIC RELEASE VERIFIED All levels

GPT-6 Astra built an interactive storm cabin | TonkaToyXL #Shorts

A quick tour of Astra’s interactive storm cabin, including the disclosed label-overlap patch.

Watch on YouTube
Short

ASTRA TEST

PUBLIC RELEASE VERIFIED All levels

GPT-6 Astra built this cabin rescue game | TonkaToyXL #Shorts

Three supplies, a snowy forest, and sixty seconds to get home. Watch Astra’s rescue-game test in short form.

Watch on YouTube
Video

ASTRA TEST

PUBLIC RELEASE VERIFIED All levels

GPT-6 Astra builds an interactive storm cabin | TonkaToyXL

Wind, snow, blackout, recovery, and controls put to the test. Five bounded checks passed; one visual patch is disclosed.

Watch on YouTube
Video

ASTRA TEST

PUBLIC RELEASE VERIFIED All levels

GPT-6 Astra builds a playable cabin rescue game | TonkaToyXL

The first game attempt passed five bounded functional checks with no code repairs. One task and one fixed seed define the result.

Watch on YouTube
Short

RECOVERY REPORT

PUBLIC RELEASE VERIFIED All levels

Luna vs Qwen vs DeepSeek: The Blank AI Games Recovered #Shorts

Blank panels became real games after a transport change. Delivery recovered, but the gameplay results still split.

Watch on YouTube
Video

RECOVERY REPORT

PUBLIC RELEASE VERIFIED All levels

We Retried the Blank AI Game Benchmarks: Luna vs Qwen vs DeepSeek

The recovery investigation separates transport failures from model performance and preserves the original locked benchmark.

Watch on YouTube
Short

MODEL DISPATCH

PUBLIC RELEASE VERIFIED All levels

OpenCode Stealth Model: What Is Omen Alpha?

A short overview of the Omen Alpha route and the evidence around its still-unconfirmed identity.

Watch on YouTube
Video

MODEL DISPATCH

PUBLIC RELEASE VERIFIED All levels

Omen Alpha Stealth Model Overview: OpenCode Go, GLM Clues, and 4 One-Shot Demos

Catalog clues and four web demos, with speculation separated from confirmed information. A research overview, not a cabin benchmark.

Watch on YouTube
Video

TOOL DISPATCH

PUBLIC RELEASE VERIFIED All levels

Hermes Just Made Local AI Setup One Click

A short dispatch on the local-AI setup workflow covered in the full episode.

Watch on YouTube
Video

TOOL DISPATCH

PUBLIC RELEASE VERIFIED All levels

Hermes Agent Makes Local AI One Click: What Nous + NVIDIA Changed

An overview of the Nous and NVIDIA local-AI setup announcement and what the workflow changes.

Watch on YouTube
Video

MODEL DISPATCH

PUBLIC RELEASE VERIFIED All levels

GPT 6 Astra is here!

The cabin’s Astra arrival dispatch. Continue to the newer hands-on builds for task-specific evidence.

Watch on YouTube
Video

ASTRA DISPATCH

PUBLIC RELEASE VERIFIED All levels

GPT-6 Astra: Real Benchmarks, Pricing, Access and Safety

The Astra launch breakdown: benchmarks, pricing, access, and safety. Follow the field guide for current primary-source specifications.

Watch on YouTube
Video

CODING TEST

PUBLIC RELEASE VERIFIED All levels

Qwen 3.8 Max vs Gemini 3.8 Flash vs Muse Spark 1.3 | One-Shot Cabin Battle

Same two visual coding prompts (Snowglobe Cabin 60, Lantern Run); one attempt per benchmark, no repair prompts, tools, web access or hints.

Watch on YouTube
Short

MODEL DISPATCH

PUBLIC RELEASE VERIFIED All levels

Meta Muse Spark 1.3: 1M Context & 88.8 Benchmarks (Wait For The Catch!)

Short companion to the Muse Spark 1.3 audit.

Watch on YouTube
Video

MODEL DISPATCH

PUBLIC RELEASE VERIFIED All levels

Meta Muse Spark 1.3: 1M Context & 88.8 Terminal-Bench (Hype or Real?)

Muse Spark 1.3 release audit; entire video produced end to end with Gemini 3.8 Flash High.

Watch on YouTube
Video

MODEL DISPATCH

PUBLIC RELEASE VERIFIED All levels

Gemini 3.8 Flash Is Here: Benchmarks, Speed, Price, and What Changed

A September 2 release audit: model facts, pricing, tools, and the independent benchmark evidence available at publication.

Watch on YouTube
Short

MODEL DISPATCH

PUBLIC RELEASE VERIFIED All levels

Gemini 3.8 Flash is officially live.

Short companion to the Gemini 3.8 Flash release audit.

Watch on YouTube
Video

MODEL DISPATCH

PUBLIC RELEASE VERIFIED All levels

Gemini 3.8 Flash is officially live.

The launch dispatch for Gemini 3.8 Flash, including what was confirmed and what still needed testing.

Watch on YouTube
Video

MODEL DISPATCH

PUBLIC RELEASE VERIFIED All levels

Qwen3.8-Max Just Hit #1 on CodeArena WebDev: What the Score Really Means

Read the September 2 CodeArena WebDev snapshot carefully: what the ranking measures and what it cannot tell you.

Watch on YouTube
Short

MODEL DISPATCH

PUBLIC RELEASE VERIFIED All levels

Qwen3.8-Max Just Hit #1 on CodeArena WebDev: What the Score Really Means

Short companion to the CodeArena WebDev snapshot.

Watch on YouTube
Video

TOOL DISPATCH

PUBLIC RELEASE VERIFIED All levels

What Is Pstack? Lauren Tan's Guide to Better AI Coding

Breakdown of Pstack, Lauren Tan's verification-first Cursor plugin: goal, build, control, evidence, retry loop.

Watch on YouTube
Short

MODEL DISPATCH

PUBLIC RELEASE VERIFIED All levels

Meta has a $5 coding plan, but is Muse Code actually worth it?

Short companion to the Muse Code plan breakdown.

Watch on YouTube
Video

TOOL DISPATCH

PUBLIC RELEASE VERIFIED All levels

Meta has a $5 coding plan. What do you actually get?

A launch-era breakdown of the Muse Code beta plan, its included features, and the limits to check before signing up.

Watch on YouTube
Short

CODING TEST

PUBLIC RELEASE VERIFIED All levels

GLM 5.3 Flash: Same Three Tests

Short companion to the GLM 5.3 Flash cabin tests.

Watch on YouTube
Video

CODING TEST

PUBLIC RELEASE VERIFIED All levels

GLM 5.3 Flash: Same Three Cabin Tests, One Model

Three frozen prompts, one run each, same Mac; untouched first pass scored 1 of 3; A* visualizer passed cold with BFS cross-check.

Watch on YouTube
Video

CODING TEST

PUBLIC RELEASE VERIFIED All levels

GPT-5.6 Luna (max) made this entire TonkaToyXL AI Test Cabin episode.

Three frozen coding prompts on the same Mac, with a procedural cabin and preserved output from the verified Luna route.

Watch on YouTube
Video

CODING TEST

PUBLIC RELEASE VERIFIED All levels

Qwen3.8 Max vs Qwen3.8 Flash: Same Three Tests, One Cabin

Same three frozen prompts on the same Mac; both went 0/3 untouched first pass, then 3/3 after one fix per test.

Watch on YouTube
Short

CODING TEST

PUBLIC RELEASE VERIFIED All levels

Qwen3.8 Flash Built This Entire Video #Shorts

Short companion to the Qwen3.8 Flash produced episode.

Watch on YouTube
Video

CODING TEST

PUBLIC RELEASE VERIFIED All levels

Qwen3.8 Flash Built This Entire Video (3 AI Coding Tests)

Episode produced by Qwen3.8 Flash including benchmark code, research, composition and edit; three frozen prompts, one run per task, 16 GB Mac mini.

Watch on YouTube
Short

CODING TEST

PUBLIC RELEASE VERIFIED All levels

OpenAI's Jalapeño Beat NVIDIA in 3 Tests, Not Everywhere #Shorts

Short companion to the Jalapeño chip test audit.

Watch on YouTube
Video

CODING TEST

PUBLIC RELEASE VERIFIED All levels

Did OpenAI's Jalapeño Chip Beat NVIDIA? Three Tests, Honest Limits

On three named OpenAI-reported InferenceX comparisons, yes; results do not prove a universal win.

Watch on YouTube
Video

TOOL DISPATCH

PUBLIC RELEASE VERIFIED All levels

Cursor vs OpenAI: What the Proposed November 12 Cutoff Means

An explanation of the reported contract cutoff proposal and its implications for Cursor users at the time of publication.

Watch on YouTube
Video

TOOL DISPATCH

PUBLIC RELEASE VERIFIED All levels

Grok Bot Templates: What They Can Actually Do

A closer look at the template workflows covered in the companion short.

Watch on YouTube
0:30

CABIN WELCOME

PUBLIC RELEASE VERIFIED All levels

Welcome to TonkaToyXL | The AI Test Cabin

A 30-second tour of Operator Tonka, the cabin crew, and the evidence-first tests that happen here.

Watch on YouTube
4:00

MODEL PREVIEW

PUBLIC RELEASE VERIFIED AI curious

Tencent Hy4 Preview: 770B Parameters, Real Value or Hype?

A source-labeled look at Tencent's large open-weight preview, its practical promise, and the questions still open.

Watch on YouTube
4:30

PHYSICAL AI

PUBLIC RELEASE VERIFIED Curious builder

This $399 Robot Could Make Physical AI Accessible

Why an affordable reinforcement-learning robot matters, and what the available evidence does and does not prove.

Watch on YouTube
2:04

INCIDENT REVIEW

PUBLIC RELEASE VERIFIED All levels

1,200 AI Agents and the Hugging Face Incident: What the Reports Say

A careful breakdown of the reported incident, the evidence the sources support, and what remains unclear.

Watch on YouTube

A strong visual still needs an honest label.

TonkaToyXL distinguishes verified product facts, vendor-reported claims, generated concepts, prototypes, historical footage, and current proof. We would rather leave out a dramatic claim than make a video less trustworthy.

See the labels and correction policy →