Comparisons

GPT-6 Astra vs Fable: A One-Shot 3D Stress Test

Direct Answer

GPT-6 Astra closes the visual-generation gap that separated earlier GPT models from Claude Fable, but Peter Gostev’s test does not produce a universal winner. Astra’s best one-shot scenes combine strong spatial coherence, clean transitions, restrained interface design, and substantial detail. Fable still has an edge in some action-heavy environments, where movement and ambient behavior appear more immediate.

The more useful discovery is that Astra’s lower reasoning levels are now viable. Low already produced scenes Peter considered stronger than the highest effort available in older GPT generations. Medium and High usually added scope and detail. Max and Ultra could add more elements, but they did not consistently improve taste, composition, or factual accuracy.

The practical verdict: treat Astra and Fable as two frontier options with different tendencies. Start Astra at Low or Medium, raise effort only when a defined rubric shows a real gap, and test both models on the workflow that matters instead of preserving last generation’s ranking.

Watch Peter’s GPT-6 Astra Gauntlet

Credits and disclosure: project footage, visual judgments, elapsed-time descriptions, and reasoning-level comparisons come from Arena AI’s video, published on 3 September 2026. Peter describes the later gauntlet as zero cherry-picking. These remain creator-run qualitative tests rather than repeated, blinded model evaluations.

What Was Actually Tested

The video contains two different kinds of evidence. The opening projects are extended builds: a multi-era London experience, a Blender-plus-Three.js D-Day scene, and an open-world game requested to contain roughly ten hours of play. Peter says some ran for hours or overnight. They show the upper end of what a long-running agent can assemble, not what every prompt returns quickly.

The later gauntlet is closer to a level playing field. Peter runs one-shot prompts at Max reasoning, generally taking around 20 to 40 minutes, and shows mistakes rather than selecting only the strongest outputs. That transparency is valuable, but a single run still mixes model ability with sampling variance.

Test groupTypical runUseful signalMain limitation
Ambitious showcasesSeveral hours or overnightLong-horizon construction across assets and codeNot one-shot and not directly comparable
Max one-shot gauntletAbout 20–40 minutesFirst-pass planning, visual synthesis, UI, motionOne sample per prompt
Reasoning ladderLow through UltraHow effort changes density and scopeRandom variation can look like an effort effect
Fable head-to-headSame prompt familyRelative visual tendenciesHarness and settings may still differ

Three Ambitious Builds Show the New Ceiling

Voxel London through the ages

The London project moves from Roman and Saxon periods through medieval, Tudor, Great Fire, and later versions of the city. The remarkable part is not the voxel style. It is the amount of state, historical framing, navigation, and playable detail held inside one experience.

This also creates a verification burden. Historical coherence and visual coherence are different. A credible production version would need source notes for dates, architecture, street plans, and transitions, plus performance checks across the full timeline.

Omaha Beach across Blender and Three.js

The D-Day scene combines assets made in Blender with a browser environment built in Three.js. Soldiers, vessels, vehicles, and environmental movement create a far more active scene than a static reconstruction. Peter is careful to say he cannot verify the military accuracy, which is exactly the right boundary.

For historical, scientific, or educational work, visual plausibility should never stand in for subject-matter review. The model can build the simulation layer; a qualified reviewer must establish whether the uniforms, vehicles, geography, sequence, and claims are defensible.

An open-world game requested to last ten hours

The game contains movement, enemies, water, sound, and a world large enough to explore. Peter had not played it for ten hours, so the duration remains a requested feature rather than a verified result. That distinction is small but important: generated scope is not tested scope.

The demo does show that long-running agents can now assemble enough systems for a meaningful game prototype. Moving from prototype to product still requires progression design, encounter variety, save behavior, performance, accessibility, balance, and sustained playtesting.

The One-Shot Gauntlet: Excellent Transitions, Visible Misses

Astra’s strongest one-shot scenes were not simply detailed. They contained transformations that required the model to connect several states: an X-ray transition, a journey into the structure of a leaf, pyramid construction over time, and movement through an impressionist painting.

The X-ray scene stood out because the transition itself felt designed rather than appended. The leaf zoom had a similar strength, turning scale into a navigable explanation. The pyramid scene was coherent, but Peter found its activity too synchronized and less alive than some Fable generations.

The mistakes remained useful. Stonehenge appeared reconstructed rather than in its present condition. A bridge near the Neuschwanstein-style castle floated. Westminster included a rotated church. These are not cosmetic edge cases when the scene represents a real place.

PromptAstra strengthObserved weaknessWhat to verify
CastleStrong overall composition and cleaner UIFloating bridgeGeometry, support, landmark references
X-ray transitionComplex state change with visual coherenceUnknown technical accuracyCorrect anatomy or mechanism
Leaf zoomHigh-quality transitions across scalePlausibility can mask scientific errorsBiology and labels
Pyramid constructionCoherent build sequenceActivity felt uniform and stagedHistorical claims and behavioral variety
Impressionist worldStyle translated into an explorable sceneSubjective fidelityComposition, interaction, accessibility
WestminsterDetail, water, and atmosphereIncorrect landmark orientationMaps, photographs, and spatial layout

Astra Versus Fable Is Now a Real Contest

Peter’s previous working assumption was that GPT-5.6 Sol trailed Fable on this class of generation. Astra changes that. In the final Westminster, White House, and dinosaur comparisons, both models produced frontier-quality work and each exposed different sharp edges.

Astra often looked stronger in transitions, spatial coherence, interface restraint, and the completeness of an explanatory sequence. Fable sometimes created richer environments with more immediate movement. In the underwater scene, Peter preferred Fable’s fish motion. In other cases, Fable’s extra dynamism came with odd geometry or less disciplined composition.

DimensionAstra tendency in the videoFable tendency in the videoDecision rule
TransitionsStrong multi-state visual sequencesCapable but prompt-dependentTest the actual transformation
Environmental motionSometimes slower to reveal activityOften more immediately dynamicMeasure behavior over time
Interface designLess overbearing than earlier GPT outputsCan be rich but occasionally denseUse a blind usability review
Real-place fidelityBeautiful with visible spatial errorsAlso vulnerable to landmark mistakesRequire reference-based QA
Overall resultGenuinely competitiveStill exceptionalRoute by workflow, not allegiance

Reasoning Levels Have Diminishing Returns

The reasoning ladder may be the most practical part of the video. Low already generated a coherent, attractive environment. Medium expanded the canvas and added detail. High and Extra High generally increased richness again. Max and Ultra produced larger or denser outputs, but did not consistently improve color, composition, or usefulness.

Peter occasionally preferred Max to Ultra. That does not prove Max is intrinsically better; it shows why one visual sample cannot isolate an effort effect from random variation. Higher effort buys more opportunity to plan, execute, and verify. It does not guarantee better taste.

EffortObserved roleGood default forEscalate when
LowAlready coherent and surprisingly capableExploration, drafts, simple scenesImportant detail or breadth is missing
MediumMore scope and environmental richnessStrong first production candidateThe rubric demands deeper execution
High / Extra HighMore elements, motion, and refinementComplex client-facing prototypesVerification or system depth still fails
MaxLarge, ambitious one-shot buildsHigh-value final candidatesOnly after cheaper levels miss clearly
UltraMaximum effort, not guaranteed tasteRare, consequential experimentsUse only with a measured reason

More Reasoning Must Earn Its Cost

OpenAI lists Astra at $10 per million input tokens and $50 per million output tokens for standard API processing. Fast mode offers up to twice the speed at twice the standard price. The video does not provide a normalized cost table for each reasoning level, so it would be misleading to assign exact dollar multipliers to Low, Medium, Max, or Ultra.

What the test makes clear is that higher effort consumes more time and tokens while quality gains flatten. The correct optimization target is not maximum reasoning. It is the lowest effort that clears the acceptance criteria with an acceptable failure rate.

Default policy: begin at Medium for new visual workflows, use Low for broad exploration, and promote to High or Max only after a scored review identifies what the cheaper run failed to deliver. Ultra should be an experiment, not a habit.

A Better Private Astra-versus-Fable Test

  1. Choose three representative tasks. Use one real-place scene, one transformation, and one interactive application from completed work.
  2. Freeze the conditions. Match the prompt, files, tools, timeout, reasoning budget, and stopping rule as closely as each product allows.
  3. Run each condition three times. This reveals whether a strong result is repeatable or merely fortunate.
  4. Review blind. Hide model names and prices while scoring the artifacts.
  5. Separate the rubric. Score factual fidelity, visual composition, movement, interaction, code structure, performance, and accessibility independently.
  6. Inspect the real app. Use browser tests, console logs, screenshots, frame-rate checks, and direct play rather than the model’s completion claim.
  7. Calculate accepted-result cost. Include token spend, wall time, failed attempts, and the human minutes needed to repair each output.

For historical and educational scenes, add a source audit. For games, verify the mechanics over time. For security work, use an authorized isolated environment: OpenAI says Astra meets its Critical cybersecurity threshold, so model capability is not permission to test systems you do not own.

Peter’s 3D Prompt Collection

The prompts behind this test family are available in Peter’s 3D Prompt Collection. The repository provides the prompts in presentation order and as structured JSON, making it a useful starting point for a repeatable visual benchmark.

A prompt alone does not reproduce the experiment. Record the model version, date, harness, tools, reasoning level, timeout, token usage, and any retries beside each output. Without that run manifest, future comparisons become visual anecdotes again.

Video Chapters

TimeTopicTimeTopic
00:00GPT-6 Astra has landed11:58Impressionist painting world
00:19Voxel London through the ages12:56Van Gogh’s house walkthrough
01:39Omaha Beach in Blender and Three.js14:48Great Fire of London and the Blitz
03:25Open-world exploration game20:04Humpback whale
05:36Max-reasoning one-shot gauntlet20:37Research-based Swiss Alps trail
07:15Neuschwanstein-style castle23:05Low through Ultra reasoning levels
09:06X-ray transition and underwater world28:51Westminster, White House, and dinosaurs
10:48Zooming into a leaf31:00Is Astra finally on par with Fable?
11:20Pyramid construction

Verdict

Peter’s test supports a meaningful update to the model hierarchy: GPT models no longer look obviously behind Claude Fable on ambitious browser-native 3D work. Astra can produce coherent worlds, sophisticated transitions, usable interfaces, and interactive systems at a level that makes the comparison prompt-dependent.

Fable remains excellent and sometimes feels more alive. Astra sometimes looks more controlled. Both can make confident spatial or factual mistakes. Max and Ultra can add richness, but Low and Medium are capable enough that sending every task to the highest effort is difficult to justify.

The right response is to reopen your evaluation, not switch camps. Put Astra and Fable against the same real tasks, start with moderate effort, review the artifacts blind, and route by accepted-result quality, cost, and risk.

Sources and Links

This article uses the primary video’s official YouTube publication date of 3 September 2026 and was researched and updated on 6 September 2026. Model access, pricing, effort controls, and product behavior can change. Recheck the linked primary sources before making a production decision.

Common questions

Is GPT-6 Astra better than Claude Fable 5.1 for 3D generation?
Peter Gostev’s tests show Astra is now genuinely competitive. Astra often produced stronger coherence, transitions, and interface restraint, while Fable sometimes created richer or more immediately dynamic scenes. Neither won every prompt.
Were all the GPT-6 Astra examples one-shot generations?
No. Voxel London, the D-Day environment, and the open-world game were ambitious multi-hour projects. The later comparison gauntlet used one-shot Max-reasoning generations that typically ran for roughly 20 to 40 minutes.
Does Ultra reasoning always produce the best result?
No. In Peter’s visual comparisons, higher effort generally added detail and scope, but Max or Ultra did not always improve composition or color. Low and Medium were already strong, and output variance can dominate small effort differences.
What does zero cherry-picking mean in this video?
It means Peter presents the one-shot outputs he received, including visible mistakes. That improves transparency, but one run per prompt is still sensitive to chance and does not replace repeated, blinded evaluation.
How should teams choose an Astra reasoning level?
Start at Low or Medium for exploration and bounded work. Increase effort when the rubric shows missing depth, verification, or complexity. Reserve Max or Ultra for high-value tasks where measured improvements justify the additional tokens and time.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call