Build & debug
Coding, tool use, and complex reasoning in one workflow.
What to check
- Does the main user flow work?
- Does it work on a narrow screen?
- Are errors handled clearly?
THE ASTRA FIELD GUIDE / SEPTEMBER 2026
GPT-6 Astra is in the cabin. Explore the documented capabilities, choose a task, and take a clear test prompt into your own session.
Vendor documentation · Checked September 5, 2026 · Model specifications ↗. Features depend on the app, tools, and access available to you.
CAPABILITY MAP
Coding, tool use, and complex reasoning in one workflow.
Research with web search and connected tools.
Document creation and long-context reasoning.
Text and image input, with text output.
Computer use, MCP, and asynchronous tool calling.
Mid-turn steering and adjustable reasoning effort.
Capability descriptions: OpenAI model card and Astra guide. Test ideas and success checks are TonkaToyXL’s suggested exercises, not published benchmark results.
THE PROMPT WORKBENCH
A useful test needs a real task, a defined output, and a way to tell if it worked.
01 / SET UP YOUR TEST
This builds a prompt in your browser. It does not send your text to a model or run a paid API call.
02 / TAKE IT FOR A TEST
Copy this into a session where you have Astra access. Save the output and compare it against the success checks.
NEW FINDINGS / SEPTEMBER 4–5
Results reported in the latest TonkaToyXL videos. Follow each one back to the test and its limits.
The published Cabin Rescue Run report records five passes on the first attempt, with no code repairs. One task, one fixed seed; the capture does not measure real-time performance.
Watch the rescue test ↗Storm Watch passed five bounded checks. One Astra-authored patch fixed a label overlap. It is an illustrative simulation, not a weather forecast or engineering model.
Watch the storm test ↗The latest comparison covers a kart game, skybridge, and Blender cabin. It preserves regressions and unavailable outputs. Estimated API-equivalent costs are not actual subscription invoices.
Inspect the full comparison ↗WHAT’S NEW
Async tools let supported applications keep independent work moving. Mid-turn steering lets you revise instructions during a task. Reasoning effort can change during a conversation while preserving the prompt cache.
See feature requirements ↗THE CABIN STANDARD
Our latest videos report bounded task results. Record the exact model, prompt, tools, effort, retries, time, and output before making a comparison. A task result is not a universal ranking.
Our existing benchmark comparison is a dated August 29 snapshot. It does not include an Astra ranking.
Inspect the benchmark snapshot →How we label evidence →