Direct Answer
GPT-6 Astra did not sweep Claude Fable 5.1 in Pat Simmons's four-build comparison. The test ended in a 2-2 tie, but Astra's reported completion cost made the result more interesting than the score. Fable produced the preferred kart racer and 3D dive-watch page. Astra produced the preferred physics simulator and website recreation. GPT-5.6 Sol generally trailed both on visual quality.
The surprise was total task economics. Pat says Astra and Fable carried the same listed token price in his comparison, yet Astra's recorded runs were often much cheaper because it completed tasks with fewer or shorter agent turns. On the Theo Jansen simulator, he reported about $6.37 for Astra versus $84 for Fable. On the site-recreation task, the figures were $71 and $223.
Watch the Four-Build Comparison
Credits and disclosure: the builds, rankings, elapsed times, and task-cost figures come from Pat Simmons's AI for Mortals video, published on 4 September 2026. The hidden model columns were revealed after ranking. This is a creator-run qualitative comparison, not a repeated independent benchmark.
How the Test Worked
Pat sent four prompts to GPT-6 Astra, Claude Fable 5.1, and GPT-5.6 Sol. Six terminal windows ran the builds in parallel, and he reviewed the applications without initially knowing which model created each column. The tasks were visual and interactive: a kart racer, a mathematically constrained walking machine, a scroll-driven product page, and a recreation of a site selected from Awwwards.
That blind reveal reduces brand bias, but it does not remove every variable. Each task appears to have one output per model, the reviewer knew the three candidates, and the criteria were not converted into a scored rubric before inspection. Agent duration, tool use, and generated volume also differed.
| What the test shows | What it does not prove | Why it matters |
|---|---|---|
| First-pass visual and interactive quality | Repeatability across multiple runs | Useful for discovering tendencies |
| Creator-observed task totals | A universal cost multiplier | Highlights execution efficiency |
| Relative performance on four prompts | Overall model intelligence | Keeps the verdict task-specific |
| Hidden identities during ranking | A fully blinded independent study | Reduces reviewer bias |
The Four-Build Scorecard
| Build | Pat's winner | Why it won | Caveat |
|---|---|---|---|
| 3D kart racer | Fable 5.1 | Better polish, environment, and play feel | One game per model |
| Theo Jansen walker | GPT-6 Astra | Cleaner controls and stronger 3D presentation | Physics not independently checked |
| 3D dive-watch page | Fable 5.1 | Smoother choreography and richer detail | Visual judgment |
| Site recreation | GPT-6 Astra | Closest recreation and strongest workflow result | Models chose different targets |
Build 1: Fable Won the Kart Racer
The first prompt asked for a polished 3D kart-racing game with responsive controls, items, opponents, barriers, and a credible track. Pat preferred Fable's result. Astra's build was playable but less polished, while Sol's version suffered from confusing controls, weak graphics, and missing environmental boundaries.
Coding strength does not automatically create art direction or game feel. A production review should test collision behavior, frame rate, input latency, opponent logic, restart states, accessibility, and whether a new player understands the controls.
Build 2: Astra Won the Walking Machine
The second build modeled a Theo Jansen-style strandbeest, a linkage system whose proportions determine whether its legs can walk. Astra earned Pat's top rank for its 3D treatment and cleaner controls. Fable offered more controls, but the interface felt busier.
Pat openly says he cannot certify the mechanism's mathematics. A convincing animation can still be mechanically wrong.
Build 3: Fable Won the Dive-Watch Page
The third prompt requested a complete scrolling product page centered on a real-time 3D dive watch. The watch needed to move from wireframe to solid, explode into components, display annotations, and reassemble as the visitor scrolled.
Fable's page was the clearest win. It combined smooth transitions, convincing lighting, detailed components, readable annotations, and a longer product narrative. Astra had attractive moments, but its watch and transitions did not reach the same level. Sol's output was less realistic and less functional.
High-fidelity product storytelling combines geometry, lighting, motion, typography, and pacing. A technically complete page can still lose when those systems do not feel like one art-directed experience.
Build 4: Astra Won the Site Recreation
For the final task, each model selected a site from Awwwards and attempted to reproduce it with Pat's public site-recreation skill. The workflow asks agents to inspect views, gather screenshots and styles, build the interface, and check their work.
Astra produced the preferred result. Pat found its target selection distinctive and its recreation impressively close in layout, interactions, transitions, and overall character. Fable also delivered a strong page. Sol failed to reproduce its selected target convincingly.
This tests more than visual generation. The agent must browse, extract evidence, follow a staged method, maintain context, write code, and compare its output with a reference. Astra's win supports a narrow but useful claim: it handled this long, tool-using workflow especially well.
Website recreation also has a rights boundary. Use it for authorized redesigns, migrations, internal prototypes, and implementation study. Replace third-party trademarks, photography, copy, and proprietary assets before publishing unless you have permission.
The Cost Surprise: Sticker Price Is Not Task Price
Matching per-token prices did not produce matching bills in Pat's runs:
| Task | Fable 5.1 | GPT-6 Astra | Interpretation |
|---|---|---|---|
| Kart racer | About $55 | About $10 | Fable won quality; Astra spent less |
| Walking machine | About $84 | About $6.37 | Astra won quality and cost |
| Site recreation | About $223 | About $71 | Astra won with a lower total |
These figures belong to Pat's setup. Total cost changes with the harness, context caching, tool calls, retry policy, reasoning setting, provider billing, and allowed run time. The video does not establish a permanent 5x to 15x advantage across every workload.
It does establish the right unit of analysis. A cheap model can become expensive through loops, retries, and repair. A higher-priced model can be economical if it reaches acceptance quickly. Track cost per accepted result, not price per million tokens alone.
A Better Model Test for Your Team
- Choose real tasks. Select completed jobs that represent your actual coding, design, research, and tool use.
- Freeze the environment. Match the prompt, files, tools, timeout, reasoning level, and stopping rule.
- Run each condition three times. One generation can be unusually good or bad.
- Review without model names. Score functionality before visual preference.
- Use a written rubric. Include correctness, visual quality, accessibility, security, maintainability, and instruction following.
- Verify the artifact. Run tests, inspect the console, exercise every control, and compare against the brief.
- Calculate accepted-result cost. Add model spend, failed runs, wall time, and human repair minutes.
The output should be a routing policy, not a single winner. One model may be your default for interactive design, another for fast implementation, and a cheaper model for drafts. Re-run the test when versions, prices, or your workflow change.
Video Chapters
| Time | Topic | Time | Topic |
|---|---|---|---|
| 00:00 | Introduction | 04:36 | Build 1: 3D kart racer |
| 00:43 | Kicking off the builds | 09:37 | Build 2: Theo Jansen walking machine |
| 00:55 | The benchmarks | 15:11 | Build 3: 3D dive-watch page |
| 02:36 | Coding benchmarks | 19:59 | Build 4: recreating an Awwwards winner |
| 02:56 | Knowledge-work benchmarks | 26:07 | Final reveal and cost |
| 26:47 | The verdict | 27:22 | Has OpenAI caught Anthropic? |
| 28:13 | What comes next |
Verdict
Pat's test supports a measured update: GPT-6 Astra is competitive with Claude Fable 5.1 on ambitious visual coding work. It won two of four builds, including the workflow that required the broadest combination of browsing, evidence extraction, code generation, and visual comparison. Fable remained stronger on the kart racer and polished product page.
The larger shift is economic. In these runs, Astra often reached a competitive result with much less recorded spend. That makes agent efficiency a first-class model characteristic, alongside intelligence, speed, and style.
OpenAI may have closed the visible capability gap in this test, but a 2-2 creator comparison is not a permanent crown. Reopen your own evaluation, measure accepted-result cost, and let repeatable evidence decide which model gets each job.
Sources and Links
- AI for Mortals: GPT-6 Astra Is HERE (Better Than Fable 5.1?)
- AI for Mortals: build links and prompts
- Pat Simmons: public site-recreation skill
- OpenAI: GPT-6 Astra
This article uses the primary video's official YouTube publication date of 4 September 2026 and was researched and published on 6 September 2026. Rankings and cost totals are attributed to the creator's recorded test. Model behavior, access, and pricing can change.