Direct Answer
GPT-6 Astra closes the visual-generation gap that separated earlier GPT models from Claude Fable, but Peter Gostev’s test does not produce a universal winner. Astra’s best one-shot scenes combine strong spatial coherence, clean transitions, restrained interface design, and substantial detail. Fable still has an edge in some action-heavy environments, where movement and ambient behavior appear more immediate.
The more useful discovery is that Astra’s lower reasoning levels are now viable. Low already produced scenes Peter considered stronger than the highest effort available in older GPT generations. Medium and High usually added scope and detail. Max and Ultra could add more elements, but they did not consistently improve taste, composition, or factual accuracy.
Watch Peter’s GPT-6 Astra Gauntlet
Credits and disclosure: project footage, visual judgments, elapsed-time descriptions, and reasoning-level comparisons come from Arena AI’s video, published on 3 September 2026. Peter describes the later gauntlet as zero cherry-picking. These remain creator-run qualitative tests rather than repeated, blinded model evaluations.
What Was Actually Tested
The video contains two different kinds of evidence. The opening projects are extended builds: a multi-era London experience, a Blender-plus-Three.js D-Day scene, and an open-world game requested to contain roughly ten hours of play. Peter says some ran for hours or overnight. They show the upper end of what a long-running agent can assemble, not what every prompt returns quickly.
The later gauntlet is closer to a level playing field. Peter runs one-shot prompts at Max reasoning, generally taking around 20 to 40 minutes, and shows mistakes rather than selecting only the strongest outputs. That transparency is valuable, but a single run still mixes model ability with sampling variance.
| Test group | Typical run | Useful signal | Main limitation |
|---|---|---|---|
| Ambitious showcases | Several hours or overnight | Long-horizon construction across assets and code | Not one-shot and not directly comparable |
| Max one-shot gauntlet | About 20–40 minutes | First-pass planning, visual synthesis, UI, motion | One sample per prompt |
| Reasoning ladder | Low through Ultra | How effort changes density and scope | Random variation can look like an effort effect |
| Fable head-to-head | Same prompt family | Relative visual tendencies | Harness and settings may still differ |
Three Ambitious Builds Show the New Ceiling
Voxel London through the ages
The London project moves from Roman and Saxon periods through medieval, Tudor, Great Fire, and later versions of the city. The remarkable part is not the voxel style. It is the amount of state, historical framing, navigation, and playable detail held inside one experience.
This also creates a verification burden. Historical coherence and visual coherence are different. A credible production version would need source notes for dates, architecture, street plans, and transitions, plus performance checks across the full timeline.
Omaha Beach across Blender and Three.js
The D-Day scene combines assets made in Blender with a browser environment built in Three.js. Soldiers, vessels, vehicles, and environmental movement create a far more active scene than a static reconstruction. Peter is careful to say he cannot verify the military accuracy, which is exactly the right boundary.
For historical, scientific, or educational work, visual plausibility should never stand in for subject-matter review. The model can build the simulation layer; a qualified reviewer must establish whether the uniforms, vehicles, geography, sequence, and claims are defensible.
An open-world game requested to last ten hours
The game contains movement, enemies, water, sound, and a world large enough to explore. Peter had not played it for ten hours, so the duration remains a requested feature rather than a verified result. That distinction is small but important: generated scope is not tested scope.
The demo does show that long-running agents can now assemble enough systems for a meaningful game prototype. Moving from prototype to product still requires progression design, encounter variety, save behavior, performance, accessibility, balance, and sustained playtesting.
The One-Shot Gauntlet: Excellent Transitions, Visible Misses
Astra’s strongest one-shot scenes were not simply detailed. They contained transformations that required the model to connect several states: an X-ray transition, a journey into the structure of a leaf, pyramid construction over time, and movement through an impressionist painting.
The X-ray scene stood out because the transition itself felt designed rather than appended. The leaf zoom had a similar strength, turning scale into a navigable explanation. The pyramid scene was coherent, but Peter found its activity too synchronized and less alive than some Fable generations.
The mistakes remained useful. Stonehenge appeared reconstructed rather than in its present condition. A bridge near the Neuschwanstein-style castle floated. Westminster included a rotated church. These are not cosmetic edge cases when the scene represents a real place.
| Prompt | Astra strength | Observed weakness | What to verify |
|---|---|---|---|
| Castle | Strong overall composition and cleaner UI | Floating bridge | Geometry, support, landmark references |
| X-ray transition | Complex state change with visual coherence | Unknown technical accuracy | Correct anatomy or mechanism |
| Leaf zoom | High-quality transitions across scale | Plausibility can mask scientific errors | Biology and labels |
| Pyramid construction | Coherent build sequence | Activity felt uniform and staged | Historical claims and behavioral variety |
| Impressionist world | Style translated into an explorable scene | Subjective fidelity | Composition, interaction, accessibility |
| Westminster | Detail, water, and atmosphere | Incorrect landmark orientation | Maps, photographs, and spatial layout |
Astra Versus Fable Is Now a Real Contest
Peter’s previous working assumption was that GPT-5.6 Sol trailed Fable on this class of generation. Astra changes that. In the final Westminster, White House, and dinosaur comparisons, both models produced frontier-quality work and each exposed different sharp edges.
Astra often looked stronger in transitions, spatial coherence, interface restraint, and the completeness of an explanatory sequence. Fable sometimes created richer environments with more immediate movement. In the underwater scene, Peter preferred Fable’s fish motion. In other cases, Fable’s extra dynamism came with odd geometry or less disciplined composition.
| Dimension | Astra tendency in the video | Fable tendency in the video | Decision rule |
|---|---|---|---|
| Transitions | Strong multi-state visual sequences | Capable but prompt-dependent | Test the actual transformation |
| Environmental motion | Sometimes slower to reveal activity | Often more immediately dynamic | Measure behavior over time |
| Interface design | Less overbearing than earlier GPT outputs | Can be rich but occasionally dense | Use a blind usability review |
| Real-place fidelity | Beautiful with visible spatial errors | Also vulnerable to landmark mistakes | Require reference-based QA |
| Overall result | Genuinely competitive | Still exceptional | Route by workflow, not allegiance |
Reasoning Levels Have Diminishing Returns
The reasoning ladder may be the most practical part of the video. Low already generated a coherent, attractive environment. Medium expanded the canvas and added detail. High and Extra High generally increased richness again. Max and Ultra produced larger or denser outputs, but did not consistently improve color, composition, or usefulness.
Peter occasionally preferred Max to Ultra. That does not prove Max is intrinsically better; it shows why one visual sample cannot isolate an effort effect from random variation. Higher effort buys more opportunity to plan, execute, and verify. It does not guarantee better taste.
| Effort | Observed role | Good default for | Escalate when |
|---|---|---|---|
| Low | Already coherent and surprisingly capable | Exploration, drafts, simple scenes | Important detail or breadth is missing |
| Medium | More scope and environmental richness | Strong first production candidate | The rubric demands deeper execution |
| High / Extra High | More elements, motion, and refinement | Complex client-facing prototypes | Verification or system depth still fails |
| Max | Large, ambitious one-shot builds | High-value final candidates | Only after cheaper levels miss clearly |
| Ultra | Maximum effort, not guaranteed taste | Rare, consequential experiments | Use only with a measured reason |
More Reasoning Must Earn Its Cost
OpenAI lists Astra at $10 per million input tokens and $50 per million output tokens for standard API processing. Fast mode offers up to twice the speed at twice the standard price. The video does not provide a normalized cost table for each reasoning level, so it would be misleading to assign exact dollar multipliers to Low, Medium, Max, or Ultra.
What the test makes clear is that higher effort consumes more time and tokens while quality gains flatten. The correct optimization target is not maximum reasoning. It is the lowest effort that clears the acceptance criteria with an acceptable failure rate.
A Better Private Astra-versus-Fable Test
- Choose three representative tasks. Use one real-place scene, one transformation, and one interactive application from completed work.
- Freeze the conditions. Match the prompt, files, tools, timeout, reasoning budget, and stopping rule as closely as each product allows.
- Run each condition three times. This reveals whether a strong result is repeatable or merely fortunate.
- Review blind. Hide model names and prices while scoring the artifacts.
- Separate the rubric. Score factual fidelity, visual composition, movement, interaction, code structure, performance, and accessibility independently.
- Inspect the real app. Use browser tests, console logs, screenshots, frame-rate checks, and direct play rather than the model’s completion claim.
- Calculate accepted-result cost. Include token spend, wall time, failed attempts, and the human minutes needed to repair each output.
For historical and educational scenes, add a source audit. For games, verify the mechanics over time. For security work, use an authorized isolated environment: OpenAI says Astra meets its Critical cybersecurity threshold, so model capability is not permission to test systems you do not own.
Peter’s 3D Prompt Collection
The prompts behind this test family are available in Peter’s 3D Prompt Collection. The repository provides the prompts in presentation order and as structured JSON, making it a useful starting point for a repeatable visual benchmark.
A prompt alone does not reproduce the experiment. Record the model version, date, harness, tools, reasoning level, timeout, token usage, and any retries beside each output. Without that run manifest, future comparisons become visual anecdotes again.
Video Chapters
| Time | Topic | Time | Topic |
|---|---|---|---|
| 00:00 | GPT-6 Astra has landed | 11:58 | Impressionist painting world |
| 00:19 | Voxel London through the ages | 12:56 | Van Gogh’s house walkthrough |
| 01:39 | Omaha Beach in Blender and Three.js | 14:48 | Great Fire of London and the Blitz |
| 03:25 | Open-world exploration game | 20:04 | Humpback whale |
| 05:36 | Max-reasoning one-shot gauntlet | 20:37 | Research-based Swiss Alps trail |
| 07:15 | Neuschwanstein-style castle | 23:05 | Low through Ultra reasoning levels |
| 09:06 | X-ray transition and underwater world | 28:51 | Westminster, White House, and dinosaurs |
| 10:48 | Zooming into a leaf | 31:00 | Is Astra finally on par with Fable? |
| 11:20 | Pyramid construction |
Verdict
Peter’s test supports a meaningful update to the model hierarchy: GPT models no longer look obviously behind Claude Fable on ambitious browser-native 3D work. Astra can produce coherent worlds, sophisticated transitions, usable interfaces, and interactive systems at a level that makes the comparison prompt-dependent.
Fable remains excellent and sometimes feels more alive. Astra sometimes looks more controlled. Both can make confident spatial or factual mistakes. Max and Ultra can add richness, but Low and Medium are capable enough that sending every task to the highest effort is difficult to justify.
The right response is to reopen your evaluation, not switch camps. Put Astra and Fable against the same real tasks, start with moderate effort, review the artifacts blind, and route by accepted-result quality, cost, and risk.
Sources and Links
- Arena AI: GPT-6 Astra | First impressions
- OpenAI: GPT-6 Astra launch page, demonstrations, benchmarks, pricing, and safety notes
- OpenAI developer documentation: GPT-6 Astra model
- Anthropic: Claude Fable 5.1
- Peter Gostev: 3D Prompt Collection
This article uses the primary video’s official YouTube publication date of 3 September 2026 and was researched and updated on 6 September 2026. Model access, pricing, effort controls, and product behavior can change. Recheck the linked primary sources before making a production decision.