THE ASTRA FIELD GUIDE / SEPTEMBER 2026

Big capabilities.
Real things to try.

GPT-6 Astra is in the cabin. Explore the documented capabilities, choose a task, and take a clear test prompt into your own session.

1.05MAPI context window
128KMaximum output tokens
Text + imagesInput · text output
5 levelsLow through max reasoning

Vendor documentation · Checked September 5, 2026 · Model specifications ↗. Features depend on the app, tools, and access available to you.

CAPABILITY MAP

Choose the work. Then test it.

01

Build & debug

Coding, tool use, and complex reasoning in one workflow.

What to check

  • Does the main user flow work?
  • Does it work on a narrow screen?
  • Are errors handled clearly?
02

Research & verify

Research with web search and connected tools.

What to check

  • Do the sources support each claim?
  • Are facts separated from inference?
  • Are source and event dates clear?
03

Create & analyze

Document creation and long-context reasoning.

What to check

  • Are the source facts preserved?
  • Do the calculations reconcile?
  • Is the result ready for its audience?
04

Inspect & explain

Text and image input, with text output.

What to check

  • Is each finding visible in the image?
  • Are assumptions clearly labeled?
  • Can someone act on the advice?
05

Work across tools

Computer use, MCP, and asynchronous tool calling.

What to check

  • Did it use the available tools?
  • Did it respect the task boundaries?
  • Can the work be checked and repeated?
06

Steer as you go

Mid-turn steering and adjustable reasoning effort.

What to check

  • Did it incorporate the correction?
  • Did completed work stay useful?
  • Did it still finish the original goal?

Capability descriptions: OpenAI model card and Astra guide. Test ideas and success checks are TonkaToyXL’s suggested exercises, not published benchmark results.

THE PROMPT WORKBENCH

Give Astra a job worth doing.

A useful test needs a real task, a defined output, and a way to tell if it worked.

01 / SET UP YOUR TEST

This builds a prompt in your browser. It does not send your text to a model or run a paid API call.

02 / TAKE IT FOR A TEST

Build & debug

Copy this into a session where you have Astra access. Save the output and compare it against the success checks.

NEW FINDINGS / SEPTEMBER 4–5

What happened in the cabin.

Results reported in the latest TonkaToyXL videos. Follow each one back to the test and its limits.

5 / 5BOUNDED FUNCTIONAL CHECKS

A rescue game, five checks.

The published Cabin Rescue Run report records five passes on the first attempt, with no code repairs. One task, one fixed seed; the capture does not measure real-time performance.

Watch the rescue test
1 patchVISUAL REFINEMENT DISCLOSED

A storm cabin, with the patch shown.

Storm Watch passed five bounded checks. One Astra-authored patch fixed a label overlap. It is an illustrative simulation, not a weather forecast or engineering model.

Watch the storm test
3 buildsONE ATTEMPT + TWO REFINEMENTS

Four models. Uneven results.

The latest comparison covers a kart game, skybridge, and Blender cabin. It preserves regressions and unavailable outputs. Estimated API-equivalent costs are not actual subscription invoices.

Inspect the full comparison

WHAT’S NEW

More control over work in progress.

Async tools let supported applications keep independent work moving. Mid-turn steering lets you revise instructions during a task. Reasoning effort can change during a conversation while preserving the prompt cache.

See feature requirements ↗

THE CABIN STANDARD

Capabilities are a starting point.

Our latest videos report bounded task results. Record the exact model, prompt, tools, effort, retries, time, and output before making a comparison. A task result is not a universal ranking.

Our existing benchmark comparison is a dated August 29 snapshot. It does not include an Astra ranking.

Inspect the benchmark snapshot →How we label evidence →