AI Model Reviews

Claude Opus 5.5 Tested: Four Builds Against Fable 5.1 and GPT-6 Astra

Launch benchmarks tell you where a model might be strong. A real build tells you whether that strength survives a browser, a codebase, Blender, missing assets, and an hour of autonomous work.

The short answer: Claude Opus 5.5 was the most balanced model in Pat Simmons' four-build test. It ranked first on the fixed brand clone and RollerCoaster Tycoon task, then second on the open-ended design clone and Blender product page. Fable 5.1 showed the best design taste in one clone. GPT-6 Astra remained clearly strongest at detailed 3D modeling. Opus was also cheaper than Fable in every reported run.

Watch the Full No-Hype Test

Source and credit: Opus 5.5: No-Hype Full Review & Testing, published by Pat Simmons on 22 September 2026. Pat published the prompts and live links at AI for Mortals.

How the Test Worked

The same three models received four demanding prompts: rebuild an Awwwards-winning site, recreate a specific yerba mate brand experience, make a playable RollerCoaster Tycoon 2-style browser game, and model an Adidas performance shoe in Blender before placing it inside an e-commerce page.

The outputs were shown without model labels. Pat ranked them visually and functionally, revealed the model names, and then compared the harness-reported time and cost. The builds ran through his multi-agent workflow rather than a normal chat window.

The open-source SimmonsBench agent fan-out harness launches parallel Claude Code sessions, gathers the completed builds, and records duration, tokens, and estimated cost. The first two tasks also used Pat's clone-app skill, which gives the agent a repeatable screenshot, implementation, and visual-QA process.

What this measures: model behavior inside one agent harness, with the same task brief and a human review of the finished artifact. It does not isolate the base model from the tools, skill instructions, website access, image generation, or cache behavior around it.

The Four-Build Scorecard

Build1st2nd3rdWhat decided it
Awwwards cloneFable 5.1Opus 5.5GPT-6 AstraVisual taste and fidelity to the chosen source
Yerba mate brandOpus 5.5Fable 5.1GPT-6 Astra3D can, page coverage, generated assets, and working interactions
RollerCoaster TycoonOpus 5.5GPT-6 AstraFable 5.1Visual resemblance, understandable controls, and playability
Adidas shoe in BlenderGPT-6 AstraOpus 5.5Fable 5.13D geometry, detail, texture, and product presentation

Across this small set, Opus 5.5 never finished last. That consistency is more useful than saying it "won" two tasks. A production default has to handle the ordinary parts of a complex job as well as the spectacular part.

Build 1: Fable Still Had the Best Design Instinct

The open-ended prompt asked each model to choose an Awwwards winner it genuinely found impressive, then rebuild it pixel for pixel. That wording tested selection and taste before implementation even began, so the models did not reproduce the same source site.

Fable 5.1 won Pat's blind ranking. Its result felt like the strongest complete design. Opus 5.5 came second and produced a close recreation of its chosen source, but it did not generate the missing imagery that would have made the page feel finished. Astra came third.

This is a useful warning for creative benchmarks: a flexible brief can reward the model that picks the easiest or most visually compatible reference. For a stricter evaluation, fix the source URL and define which assets may be reused, generated, or replaced.

Build 2: Opus Won the Fixed Brand Clone

The second test removed that freedom. All three models had to recreate the same yerba mate brand website, including a rotating 3D can, animated sections, product imagery, and a small jumping game.

Opus 5.5 won because it delivered the broadest finished experience. Its can texture and rotation were convincing, the page included generated labels and product shots, and the game worked. Fable's 3D work was strong, but missing or broken images weakened the page. Astra completed the task quickly and cheaply, yet its overall recreation ranked third.

This result matters more than a screenshot contest. Complex product pages fail through accumulation: one missing section, one broken image path, one inert interaction, and one unconvincing 3D asset can turn a technically valid build into something a team would not ship.

Build 3: Opus Made the Most Usable Game

For the game test, each model received a detailed request for a playable browser clone of RollerCoaster Tycoon 2. Opus 5.5 produced the closest visual interpretation and the clearest controls. Astra's version was understandable but more generic. Fable's result was difficult to read and use.

Opus winning here does not mean it can replace a game team. It means the model successfully coordinated interface, simulation rules, state, controls, and recognizable visual language in one autonomous pass. The next evaluation should test deeper behavior: track placement errors, economy balance, save and reload, long-session stability, and whether another developer can extend the code.

Build 4: Astra Kept the 3D Crown

The final prompt combined two disciplines: model the Adidas Adizero Adios Pro Evo 3 in Blender, create exploded components and annotations, and embed the result in a polished product page.

GPT-6 Astra won by a wide margin on the actual 3D object. Its geometry, detail, and texture made the shoe look much closer to the real product. Its weakness was web presentation: the embedded model and page behavior were less polished than the artifact itself.

Opus 5.5 finished second. The shoe was less convincing, but the overall site integration was cleaner. Fable 5.1 ranked third. The split suggests a practical two-stage workflow: use the strongest 3D model for asset creation, then route site assembly and interaction polish to the model that performs better in your web stack.

Reported Cost and Runtime

BuildOpus 5.5Fable 5.1GPT-6 Astra
Awwwards clone$47 / 76 min$177 / 75 min$144 / 60 min
Yerba mate brand$58 / 74 min$107 / 45 min$33 / 39 min
RollerCoaster Tycoon$53 reported$84 reported$5 reported
Adidas shoe$33 / 91 min$65 / 68 min$13 / about half Opus time

These are the creator's harness-reported figures, not guaranteed prices for reproducing the builds. The first clone reported very large input-token totals: about 139 million for Opus, 127 million for Fable, and 99 million for Astra. Repeated tool context and cache behavior can make those counts look unlike a normal prompt-and-response bill.

Two conclusions are still reasonable. First, Opus was cheaper than Fable on every task shown. Second, Astra was extraordinarily cost-effective on the game and 3D tests even when it did not win the overall page ranking. The right comparison is cost per accepted artifact, not cost per million tokens or cost per run.

Benchmark Claims Need the Same Discipline

Anthropic reports Opus 5.5 at 66.4% on Terminal-Bench 4.0, 54.4% on FrontierCode 1.1, 57.8% on CursorBench 4.0, and 1846 Elo on GDPval-AA 2.1. It also says typical workloads cost 40% less than Opus 5 and output is more than 30% faster.

Those numbers establish a credible starting point. They do not prove superiority on CAD, design taste, long-horizon coding, or every computer-use workflow. Pat specifically notes tests missing from Anthropic's launch table and compares some outside results separately. Different harnesses, effort levels, safety policies, and tool access make a clean cross-vendor ranking difficult.

The four builds add practical evidence, but they have their own limits: one run per model, one reviewer, subjective ordering, and tasks that mix model intelligence with tools and skills. A stronger internal evaluation would repeat each task at least three times, blind multiple reviewers, and score defined acceptance criteria before revealing cost.

Which Model Should You Use?

WorkloadStart withWhy
Mixed web builds with many moving partsClaude Opus 5.5Most consistent result across the four tests
Open-ended visual directionOpus 5.5 and Fable 5.1 side by sideFable showed stronger taste in the first clone
Blender and detailed 3D assetsGPT-6 AstraClear lead on the shoe model and texture
Cheap exploratory prototypesGPT-6 Astra, then compare OpusVery low reported cost on the game and 3D tasks
Long, difficult coding workOpus 5.5 vs Fable 5.1 on your own suiteThe four creative builds do not settle long-horizon reliability

The most useful outcome is a routing policy, not a winner. Keep a frozen set of real tasks, compare accepted-result rate and reviewer effort, and send each workload to the least expensive model that repeatedly clears your bar.

Video Chapters

TimeChapterTimeChapter
00:00Intro00:36Kicking off the builds
00:55Pricing04:46Benchmarks
06:55Tests missing from Anthropic's table09:10Awwwards clone
17:47Yerba mate brand clone24:48RollerCoaster Tycoon clone
29:15Adidas shoe in Blender33:58Final thoughts

Sources and Tools

Common questions

Did Claude Opus 5.5 beat Fable 5.1 and GPT-6 Astra?
In Pat Simmons' four creator tests, Opus 5.5 ranked first twice and second twice. Fable 5.1 won the open-ended Awwwards clone, while GPT-6 Astra won the Blender shoe model. This is a useful directional result, not a universal model ranking.
Which model was best for 3D work?
GPT-6 Astra produced the strongest Adidas shoe model and textures in this test. Opus 5.5 integrated its 3D result into the website more cleanly, but Astra led on the actual 3D artifact.
Was Opus 5.5 cheaper than Fable 5.1?
Yes in all four reported runs. The savings varied by task. The largest caveat is that one run per task cannot establish a stable average, and these figures include the behavior of the particular prompts, tools, cache, and harness.
Was the ranking blind?
The presenter reviewed the three outputs without showing the model names first, then revealed the models and examined reported cost and runtime. The ranking still reflects one human reviewer's taste and task priorities.
Can I reproduce the test setup?
Pat Simmons published the prompts and live builds through AI for Mortals, an open-source agent fan-out harness on GitHub, and the website-cloning skill used in the first tests. Reproduction still requires compatible tools, model access, and careful permission review.
Should I replace Fable 5.1 with Opus 5.5?
Test both on your own accepted-result criteria. Opus 5.5 is the stronger default from these four runs, while Fable 5.1 may still justify its cost on long, difficult coding work where its higher ceiling produces fewer failures.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call