Comparisons

Claude Opus 5 Benchmarked: What $400 of Simulations Actually Proves

Direct Answer

Claude Opus 5 looks exceptionally strong at turning open-ended prompts into polished, interactive prototypes. Nick Saraev's creator-reported $400 test produced a 3D gallery, physics-style simulations, games, visualizations, and editing tools with unusually good first-pass presentation. Alex Finn reached a similar conclusion in a separate Opus 5 versus Fable 5 comparison.

The useful conclusion is narrower than "Opus 5 can build anything." These videos show high prototype fluency, strong visual composition, broad interaction design, and promising cost efficiency. They do not independently validate physics, historical data, production architecture, accessibility, security, maintainability, or repeatability.

Main video credit: developer and AI educator Nick Saraev. Follow Nick on X. The video title and creator identity were verified through YouTube's public metadata.

JQ AI SYSTEMS take: Opus 5 has earned a serious evaluation lane. Promote it by cost per accepted result, not by how cinematic one selected demo looks.

Evidence and Credits

This article separates three evidence layers. Nick Saraev's and Alex Finn's videos are creator tests. Anthropic's launch post and documentation provide the official pricing, benchmark methodology notes, effort behavior, and safety claims. The recommendations and verification framework below are JQ AI SYSTEMS analysis.

Source Evidence type What it supports
Nick Saraev's $400 Opus 5 test Creator test Selected prototypes, creator-observed costs, and interpretation of launch charts.
Alex Finn's Opus 5 vs Fable 5 test Creator test Five informal head-to-head tasks and operator preferences.
Anthropic launch post Vendor official Pricing, availability, evaluations, methodology notes, safeguards, and alignment claims.
Anthropic effort guide Vendor official Low-to-max effort behavior and token-efficiency controls.
Every field review and migration guide Companion analysis Why a strong model can still behave badly inside mature skills and harnesses.

The Simulation Scorecard

Nick's demos are impressive because they cover several interaction families instead of one landing page. The table below keeps the useful evidence while making the missing verification explicit.

Build What the demo demonstrates What still needs testing
3D gallery Scene composition, navigation, hover states, audio cues, and artwork cataloging. Performance on low-end devices, keyboard navigation, asset rights, and content management.
Orbit launch game Interactive controls, educational framing, visual feedback, and orbital concepts. Equations, units, numerical stability, and expected trajectories against a reference solver.
Cloth and wrecking-ball simulations Real-time visual response, user controls, collision effects, and satisfying feedback. Conservation behavior, collision correctness, timestep sensitivity, and reproducible physical outcomes.
Fractal and falling-sand systems Shader-style visuals, responsive controls, procedural behavior, and animation. Frame rate, memory growth, edge cases, device compatibility, and deterministic rules.
Predator-prey ecosystem Stateful entities, emergent-looking motion, parameter controls, and live statistics. Whether the population model implements the claimed biology rather than plausible animation.
Quadcopter and double pendulum Control mapping, motion, game feel, and a coherent interactive loop. Flight dynamics, pendulum equations, reset behavior, browser consistency, and input latency.
Pixel editor Tool selection, undo, erasing, layers, frames, onion skinning, and canvas resizing. File import/export, destructive resize behavior, large canvases, persistence, and accessibility.
Sneaker configurator Part selection, material and color controls, camera movement, and render export. Accurate product geometry, realistic materials, SKU logic, pricing, and production assets.
Population map Animated storytelling, time controls, geographic navigation, and comparative charts. Source provenance, historical boundaries, population definitions, missing years, and factual accuracy.

What the Demos Prove

First, Opus 5 has strong prototype breadth. It can combine controls, state, motion, visual hierarchy, and explanatory text in one pass. That is valuable for founders and product teams because a working artifact creates a better conversation than a static specification.

Second, the model appears comfortable generating visual material from code instead of depending entirely on downloaded assets. Nick highlights SVG and procedural elements in several builds. This can reduce setup friction and make a prototype more self-contained, although generated approximations are not substitutes for licensed production assets.

Third, the work supports Anthropic's claim that Opus 5 improved at visual output and frontend replication. Anthropic's launch material includes interactive wind-tunnel and cell examples, while early-access partners specifically report gains in animations, games, 3D work, browser verification, and mobile layout checks.

What the Demos Do Not Prove

A simulation can look physically convincing while implementing the wrong model. Nick says this directly during the cloth example: he cannot verify whether the physics are accurate. The same caution applies to orbital mechanics, bridges, pendulums, ecosystems, and population data.

A selected first-pass result also does not establish reliability. We do not have the complete prompt set, every failed run, model settings, tool permissions, dependency versions, selection criteria, or a machine-readable result bundle. The video is excellent product evidence and weak scientific evidence.

Production readiness needs a different test: responsive behavior, accessibility, error states, persistence, authentication, security boundaries, observability, deployment, load, maintenance, and deterministic acceptance checks. A strong prototype may shorten that journey, but it does not complete it.

Practical rule: judge a visual demo twice. First ask, "Does this communicate the product?" Then ask, "Can the underlying claims survive a test harness?"

Official Benchmark Context

Anthropic says Opus 5 is state of the art on coding and knowledge-work evaluations including Frontier-Bench and GDPval-AA. On CursorBench 3.2 at max effort, Anthropic reports performance within 0.5 percentage points of Fable 5's peak at half the cost per task. It also reports a score three times higher than the next-best model on ARC-AGI 3, around 1.5 times the next-best pass rate at the same cost on Zapier AutomationBench, and stronger cost-performance on OSWorld 2.0.

Those results are more controlled than a YouTube build, but they are still launch evaluations selected and reported by the vendor or partners. Anthropic does publish useful methodology detail: its Frontier-Bench chart used the mini-SWE-agent harness, a GKE backend, and the mean reward across five attempts per task, with Opus 4.8 fallback on safety-classifier refusals for Opus 5 and Fable 5.

That final detail matters. If a request falls back to another model, record the model that actually completed the work. Otherwise a team can accidentally credit Opus 5 for an Opus 4.8 result or compare unlike safety behavior.

Cost and Effort Curves

Anthropic lists Opus 5 at $5 per million input tokens and $25 per million output tokens, the same base price as Opus 4.8. Fable 5 has a higher base token price, but the useful metric is not list price. It is cost per accepted result.

In Nick's selected Airlines website example, the displayed cost was $0.69 for Opus 5 and $0.94 for Fable 5, with Nick preferring the Opus result. Alex reports an Opus roller-coaster build at $0.82 and roughly 32,000 tokens, with Fable costing about 50 percent more in his setup. These are informative snapshots, not stable price ratios.

Effort Start here for Promotion rule
Low Small edits, narrow analysis, fast review, and high-volume sub-tasks. Keep it if the acceptance rate matches higher effort.
Medium Bounded implementation, research, business workflows, and routine tool use. A strong default candidate when quality holds and latency matters.
High Complex work that benefits from deeper reasoning; this is the API default. Compare against medium before accepting the added spend.
Xhigh Difficult coding and long-running agentic work. Anthropic recommends it as a starting point for hard agentic tasks; confirm with your eval.
Max Rare frontier tasks where capability matters more than token spend. Use only when repeated tests show a material gain over xhigh.

Count input, output, cache writes and reads, tool calls, subagents, retries, repair prompts, elapsed time, and human review. A $0.69 prototype that needs four hours of correction can be more expensive than a $2 run that passes immediately.

Alex Finn's Opus 5 vs Fable 5 Comparison

Additional video credit: creator and software founder Alex Finn. Follow Alex on X.

Alex compares five informal tasks: a 3D roller coaster, an Apple-style page reconstruction, a document scavenger hunt, open-source bug fixing, and a bridge simulator. He gives Opus 5 the overall win on visual quality and reported spend. Fable holds slightly more weight in the bridge test, while a classifier prevents it from completing the document task.

His operational verdict is more nuanced than the headline. He finds Opus 5 verbose, prone to doing adjacent work, and less pleasant for planning conversations. He still prefers Fable for large planning and back-and-forth work, GPT-5.6 for a daily driver, and Opus 5 for difficult execution.

One spoken total-cost figure in the supplied transcript is internally inconsistent with the individual examples, so it is intentionally not repeated here. The per-test costs and relative descriptions are the defensible evidence available from the walkthrough.

Alignment Is Not a Guarantee

Nick highlights Anthropic's automated behavioral-audit score of 2.3 for overall misaligned behavior, lower than the recent Claude models shown in the launch chart. Anthropic says Opus 5 follows Claude's Constitution better, shows lower deceptive behavior, is harder to trick into misuse, and is less likely to take reckless, hard-to-reverse actions.

That is encouraging evidence from one evaluation family. It does not make autonomous deployment risk-free. A model can score well on behavioral audits and still make an ordinary software mistake, misuse a broad permission, act on stale data, or optimize the wrong target.

Keep sandboxing, least-privilege credentials, reversible actions, human approval for external side effects, deterministic tests, audit logs, and production monitoring. Alignment scores inform the control design; they do not replace it.

A Better Test Protocol

  1. Choose a real task. Use a repeated workflow with business value, not a prompt designed only to produce spectacle.
  2. Freeze the environment. Keep the harness, repository, dependencies, tools, permissions, timeout, and machine class constant.
  3. Use one task contract. Give each model the same inputs, scope, definition of done, and approval boundaries.
  4. Define objective checks. Add tests, reference data, performance budgets, accessibility checks, or numerical tolerances before running.
  5. Run at least three attempts. One result measures a sample. Several runs begin to measure reliability and variance.
  6. Blind subjective review. Hide the model name when scoring design, writing, or usability.
  7. Track the full cost. Record tokens, cache behavior, tools, subagents, retries, repair prompts, elapsed time, and review minutes.
  8. Run an effort sweep. Test low, medium, high, and xhigh before deciding that more reasoning is better.
  9. Publish failures too. Keep every run in the result set rather than selecting only the strongest artifact.
  10. Promote by acceptance rate. Route production work to the model with the best quality, reliability, and total cost for that task class.
Evaluate this model on: [REAL TASK]

Environment
- Harness:
- Tools and permissions:
- Model and effort:
- Timeout and budget:

Definition of done
- Required behavior:
- Required files or deliverables:
- Tests and numerical tolerances:
- Accessibility and performance checks:
- Actions requiring approval:

Record
- Accepted on first run: yes/no
- Defects:
- Repair prompts:
- Input/output/cache tokens:
- Tool calls and subagents:
- Elapsed time:
- Human review minutes:
- Final accepted cost:

When to Use Opus 5

Work type Recommended lane Verification
Interactive prototype or visual simulation Opus 5 at medium or high, escalating only when needed. Browser tests, mobile checks, performance budget, and domain-specific correctness tests.
Large implementation or difficult debugging Opus 5 at xhigh as an evaluation starting point. Repository tests, regression suite, code review, and bounded permissions.
Ambiguous strategy or creative direction Compare Opus 5 with Fable; operator preference may matter. Decision rubric, evidence trail, and human ownership of the final direction.
Routine daily work Test lower-effort Opus, Sonnet, Sol, or another cheaper worker. Latency, accepted-task cost, and repair rate over a representative week.
Physics, science, finance, legal, or historical claims Use Opus 5 as an implementation and analysis assistant, not the authority. Reference data, domain tools, citations, independent calculation, and expert review where stakes require it.

Video Chapters

Time Nick Saraev's main video
00:00Opus 5 launch and 3D gallery demo
01:38Orbit launch game
02:30Cloth simulation
02:55Fractal generator
03:28Falling-sand cellular simulation
04:02Predator-prey ecosystem
04:32Quadcopter flight simulation
04:55Double pendulum
05:14Pixel editor
05:52Sneaker configurator
06:40Animated population map
07:14Wrecking-ball simulator
08:20Cost comparison with Fable 5
09:05Benchmark breakdown
11:22Performance-per-dollar effort curves
13:40Alignment scores
14:34What the release means
Time Alex Finn's comparison
00:00Opus 5 verdict
01:40Five Opus 5 versus Fable 5 tests
06:19Personality, limits, and harness concerns
09:45When Alex uses Opus, Fable, and GPT-5.6

Bottom Line

Opus 5 appears to be a genuine step forward for interactive coding, visual output, difficult implementation, and professional knowledge work. Nick Saraev's prototypes make that progress tangible, while Alex Finn's comparison adds useful evidence that the model can deliver strong artifacts at lower reported spend than Fable in selected tasks.

The strongest takeaway is not that every demo is production-ready or scientifically correct. It is that teams can now reach a richer prototype sooner, then spend their scarce human attention on verification, product judgment, and the parts of the system where errors matter.

Start with a real task, test multiple runs and effort levels, count repairs, verify the underlying claims, and route by accepted result. That is how a compelling launch demo becomes a dependable operating decision.

Sources

Common questions

Is Claude Opus 5 better than Claude Fable 5?
It depends on the task and evaluation method. Nick Saraev and Alex Finn preferred Opus 5 on several visual builds and reported lower spend. Anthropic reports near-Fable performance on selected coding evaluations at roughly half the cost per task. Neither source proves that Opus 5 wins every planning, coding, or production workflow.
What did Nick Saraev build with Claude Opus 5?
The video shows a 3D gallery, orbital launch game, cloth simulation, fractal generator, falling-sand simulation, predator-prey ecosystem, quadcopter game, double pendulum, pixel editor, sneaker configurator, animated population map, and wrecking-ball simulator.
Does a working simulation prove the physics are accurate?
No. A visually plausible cloth, orbit, bridge, pendulum, or collision demo proves interface generation and prototype fluency. Physical correctness requires reference equations, invariant checks, controlled test cases, numerical tolerances, and often domain review.
How much does Claude Opus 5 cost?
Anthropic lists standard API pricing at $5 per million input tokens and $25 per million output tokens. The total cost of a result also depends on effort, cache behavior, tools, subagents, retries, repair prompts, and human review.
Which Claude Opus 5 effort level should I use?
Start with the lowest effort that passes your acceptance tests. Low and medium suit bounded work, high is the default, xhigh is a reasonable starting point for difficult coding and agentic tasks, and max should be reserved for cases where repeated evaluation shows a real gain.
Is Opus 5 safer because its alignment score is lower?
Anthropic reports a lower rate of measured misaligned behavior in its automated audit, which is encouraging. It is not a guarantee that every output or autonomous action will be safe. Permissions, sandboxing, tests, approvals, and monitoring still matter.
How should I compare Opus 5 with another model?
Use the same task, harness, tools, effort policy, timeout, budget, and acceptance tests. Run each model several times, blind the human reviewer where possible, record repair prompts and review time, and compare cost per accepted result rather than the prettiest first output.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call