Launch benchmarks tell you where a model might be strong. A real build tells you whether that strength survives a browser, a codebase, Blender, missing assets, and an hour of autonomous work.
Watch the Full No-Hype Test
Source and credit: Opus 5.5: No-Hype Full Review & Testing, published by Pat Simmons on 22 September 2026. Pat published the prompts and live links at AI for Mortals.
How the Test Worked
The same three models received four demanding prompts: rebuild an Awwwards-winning site, recreate a specific yerba mate brand experience, make a playable RollerCoaster Tycoon 2-style browser game, and model an Adidas performance shoe in Blender before placing it inside an e-commerce page.
The outputs were shown without model labels. Pat ranked them visually and functionally, revealed the model names, and then compared the harness-reported time and cost. The builds ran through his multi-agent workflow rather than a normal chat window.
The open-source SimmonsBench agent fan-out harness launches parallel Claude Code sessions, gathers the completed builds, and records duration, tokens, and estimated cost. The first two tasks also used Pat's clone-app skill, which gives the agent a repeatable screenshot, implementation, and visual-QA process.
The Four-Build Scorecard
| Build | 1st | 2nd | 3rd | What decided it |
|---|---|---|---|---|
| Awwwards clone | Fable 5.1 | Opus 5.5 | GPT-6 Astra | Visual taste and fidelity to the chosen source |
| Yerba mate brand | Opus 5.5 | Fable 5.1 | GPT-6 Astra | 3D can, page coverage, generated assets, and working interactions |
| RollerCoaster Tycoon | Opus 5.5 | GPT-6 Astra | Fable 5.1 | Visual resemblance, understandable controls, and playability |
| Adidas shoe in Blender | GPT-6 Astra | Opus 5.5 | Fable 5.1 | 3D geometry, detail, texture, and product presentation |
Across this small set, Opus 5.5 never finished last. That consistency is more useful than saying it "won" two tasks. A production default has to handle the ordinary parts of a complex job as well as the spectacular part.
Build 1: Fable Still Had the Best Design Instinct
The open-ended prompt asked each model to choose an Awwwards winner it genuinely found impressive, then rebuild it pixel for pixel. That wording tested selection and taste before implementation even began, so the models did not reproduce the same source site.
Fable 5.1 won Pat's blind ranking. Its result felt like the strongest complete design. Opus 5.5 came second and produced a close recreation of its chosen source, but it did not generate the missing imagery that would have made the page feel finished. Astra came third.
This is a useful warning for creative benchmarks: a flexible brief can reward the model that picks the easiest or most visually compatible reference. For a stricter evaluation, fix the source URL and define which assets may be reused, generated, or replaced.
Build 2: Opus Won the Fixed Brand Clone
The second test removed that freedom. All three models had to recreate the same yerba mate brand website, including a rotating 3D can, animated sections, product imagery, and a small jumping game.
Opus 5.5 won because it delivered the broadest finished experience. Its can texture and rotation were convincing, the page included generated labels and product shots, and the game worked. Fable's 3D work was strong, but missing or broken images weakened the page. Astra completed the task quickly and cheaply, yet its overall recreation ranked third.
This result matters more than a screenshot contest. Complex product pages fail through accumulation: one missing section, one broken image path, one inert interaction, and one unconvincing 3D asset can turn a technically valid build into something a team would not ship.
Build 3: Opus Made the Most Usable Game
For the game test, each model received a detailed request for a playable browser clone of RollerCoaster Tycoon 2. Opus 5.5 produced the closest visual interpretation and the clearest controls. Astra's version was understandable but more generic. Fable's result was difficult to read and use.
Opus winning here does not mean it can replace a game team. It means the model successfully coordinated interface, simulation rules, state, controls, and recognizable visual language in one autonomous pass. The next evaluation should test deeper behavior: track placement errors, economy balance, save and reload, long-session stability, and whether another developer can extend the code.
Build 4: Astra Kept the 3D Crown
The final prompt combined two disciplines: model the Adidas Adizero Adios Pro Evo 3 in Blender, create exploded components and annotations, and embed the result in a polished product page.
GPT-6 Astra won by a wide margin on the actual 3D object. Its geometry, detail, and texture made the shoe look much closer to the real product. Its weakness was web presentation: the embedded model and page behavior were less polished than the artifact itself.
Opus 5.5 finished second. The shoe was less convincing, but the overall site integration was cleaner. Fable 5.1 ranked third. The split suggests a practical two-stage workflow: use the strongest 3D model for asset creation, then route site assembly and interaction polish to the model that performs better in your web stack.
Reported Cost and Runtime
| Build | Opus 5.5 | Fable 5.1 | GPT-6 Astra |
|---|---|---|---|
| Awwwards clone | $47 / 76 min | $177 / 75 min | $144 / 60 min |
| Yerba mate brand | $58 / 74 min | $107 / 45 min | $33 / 39 min |
| RollerCoaster Tycoon | $53 reported | $84 reported | $5 reported |
| Adidas shoe | $33 / 91 min | $65 / 68 min | $13 / about half Opus time |
These are the creator's harness-reported figures, not guaranteed prices for reproducing the builds. The first clone reported very large input-token totals: about 139 million for Opus, 127 million for Fable, and 99 million for Astra. Repeated tool context and cache behavior can make those counts look unlike a normal prompt-and-response bill.
Two conclusions are still reasonable. First, Opus was cheaper than Fable on every task shown. Second, Astra was extraordinarily cost-effective on the game and 3D tests even when it did not win the overall page ranking. The right comparison is cost per accepted artifact, not cost per million tokens or cost per run.
Benchmark Claims Need the Same Discipline
Anthropic reports Opus 5.5 at 66.4% on Terminal-Bench 4.0, 54.4% on FrontierCode 1.1, 57.8% on CursorBench 4.0, and 1846 Elo on GDPval-AA 2.1. It also says typical workloads cost 40% less than Opus 5 and output is more than 30% faster.
Those numbers establish a credible starting point. They do not prove superiority on CAD, design taste, long-horizon coding, or every computer-use workflow. Pat specifically notes tests missing from Anthropic's launch table and compares some outside results separately. Different harnesses, effort levels, safety policies, and tool access make a clean cross-vendor ranking difficult.
The four builds add practical evidence, but they have their own limits: one run per model, one reviewer, subjective ordering, and tasks that mix model intelligence with tools and skills. A stronger internal evaluation would repeat each task at least three times, blind multiple reviewers, and score defined acceptance criteria before revealing cost.
Which Model Should You Use?
| Workload | Start with | Why |
|---|---|---|
| Mixed web builds with many moving parts | Claude Opus 5.5 | Most consistent result across the four tests |
| Open-ended visual direction | Opus 5.5 and Fable 5.1 side by side | Fable showed stronger taste in the first clone |
| Blender and detailed 3D assets | GPT-6 Astra | Clear lead on the shoe model and texture |
| Cheap exploratory prototypes | GPT-6 Astra, then compare Opus | Very low reported cost on the game and 3D tasks |
| Long, difficult coding work | Opus 5.5 vs Fable 5.1 on your own suite | The four creative builds do not settle long-horizon reliability |
The most useful outcome is a routing policy, not a winner. Keep a frozen set of real tasks, compare accepted-result rate and reviewer effort, and send each workload to the least expensive model that repeatedly clears your bar.
Video Chapters
| Time | Chapter | Time | Chapter |
|---|---|---|---|
| 00:00 | Intro | 00:36 | Kicking off the builds |
| 00:55 | Pricing | 04:46 | Benchmarks |
| 06:55 | Tests missing from Anthropic's table | 09:10 | Awwwards clone |
| 17:47 | Yerba mate brand clone | 24:48 | RollerCoaster Tycoon clone |
| 29:15 | Adidas shoe in Blender | 33:58 | Final thoughts |
Sources and Tools
- Pat Simmons: Opus 5.5 - No-Hype Full Review & Testing
- AI for Mortals: all prompts and live build links
- GitHub: SimmonsBench agent fan-out harness
- GitHub: clone-app skill used for website recreation
- Anthropic: Introducing Claude Opus 5.5
- Anthropic: Claude Opus 5.5 System Card
- Artificial Analysis: independent model comparisons