Comparisons

Claude Opus 5 vs Fable 5: Nine Real Workflow Tests, Cost, and Verification

Direct Answer

Claude Opus 5 is not a simple Fable 5 replacement. Nate Herk's nine-workflow comparison found a more useful split: Opus 5 was excellent when correctness, instruction following, and persistent verification mattered, while Fable 5 was usually faster and often produced the stronger visual or creative artifact.

Opus 5 also costs half as much per API token. That did not make every Opus run cheaper. In several experiments it generated so many more tokens, tool calls, or verification passes that the completed task cost more than Fable. The practical metric is therefore cost per accepted result, not price per token.

JQ AI SYSTEMS take: use Opus 5 as the default high-end worker for bounded tasks with clear tests. Bring in Fable 5 when the job needs broad orchestration, creative judgment, or a fast ambitious first pass. Promote either model only after it passes your own workflow.

Watch the Test

Video credit: Nate Herk | AI Automation. Follow Nate Herk on X. YouTube's public metadata lists the original title as I Tested Opus 5 vs. Fable 5. What You Need to Know.

What Nate Actually Tested

The first experiment asked both models in Claude chat to generate an Excalidraw explanation of semantic search. There were no skills, saved personal context, or verification tools. Nate preferred Opus 5's more organized educational layout, but correctly described that judgment as subjective.

The remaining work mostly ran in Claude Code. That matters. These results compare Opus 5 and Fable 5 inside the same agent harness, with access to files, tools, skills, and long-running loops. They do not isolate model intelligence from harness behavior.

The suite covered nine workflow families: an explanatory diagram, two codebase bug hunts, an announcement video, a landing page, a LinkedIn post and carousel, audience research, a video outline and slide deck, browser-based computer use, and a structural simulator.

Nine-Test Scorecard

Workflow Reported run data Useful verdict
1. Semantic-search diagram No cost shown; raw Claude chat Opus was more organized; Fable was more visual. Subjective and unverified.
2. Codebase bug hunts Run A: Fable $5.30 / 11 min; Opus $4.22 / 13 min. Run B: Fable $8.73 / 12 min; Opus $6.50 / about 20 min. Split first run. Opus clearly won the second with four of four checks and a 93/95 technical score.
3. AIS Live announcement video Fable about $7 / 7:49; Opus $11.19 / 40 min Fable was faster. Opus delivered both vertical and landscape versions. Fable used outdated event details, exposing a context and verification failure.
4. Program landing page Fable $20.50 / 22 min; Opus $35.83 / almost 1 hour Very similar branded pages. Nate gave Fable a slight creative edge and it was substantially faster and cheaper.
5. LinkedIn post and carousel Fable $6.17 / about 7.5 min; Opus $8.22 / longer Writing was similar; Nate preferred Fable's carousel. He also noted that both premium models were probably overkill.
6. Audience research report Fable $10.60 / about 11 min; Opus $8.34 / 20 min Outputs were broadly similar. Opus was cheaper; Fable was faster. Source coverage still needs an audit.
7. Video outline and slide deck Fable about $41 / 8 min; Opus about $33 / 75 min Fable cost more but produced the deck Nate preferred and finished far faster.
8. Computer-use Snake game Opus run A about $18 / 113 min; Opus run B about $10 / 70 min Not an Opus versus Fable test. Both were accidentally Opus, revealing large run-to-run variance from the same prompt.
9. Structural simulator Fable $73 / 7 min; Opus $112 / 146 min Nate preferred Opus's usability and detail. Neither output was validated by a structural engineer, so treat both as prototypes.

All prices, durations, scores, and preferences above are creator-reported from the video and attached transcript. They are not independent benchmarks. Some displayed totals were not spoken aloud, so this article does not infer them from the screen.

The Strongest Evidence Was Not the Prettiest Demo

The second bug hunt is the most decision-useful result because it had explicit expected behavior and a separate Codex review. Opus passed four of four checks and scored 93 out of 95. Fable passed two of four, left an issue unresolved, and scored 66 out of 95. Opus also cost less in that run, although it took longer.

The first bug hunt was much closer. Both models effectively solved the problem, and the reviewer preferred Fable's cleaner one-line upstream fix. Taken together, the two runs do not prove that Opus always codes better. They do show why repository tests and independent review produce stronger evidence than visual preference.

The structural simulator was impressive, but its appearance is not proof of engineering correctness. Nate explicitly said he was not a structural engineer. Before using such software for a real decision, a team would need reference equations, units, known test cases, tolerance checks, stress scenarios, and expert review.

Why Cheaper Tokens Can Produce a More Expensive Task

Anthropic lists Opus 5 at $5 per million input tokens and $25 per million output tokens. Fable 5 is $10 and $50. The per-token price really is half. Yet Opus cost more on the landing page, LinkedIn package, announcement video, and simulator because it ran longer or generated more work.

Nate's unequal aggregate included ten Opus sessions and eight Fable sessions. Opus produced about 2 million output tokens versus roughly 832,000 for Fable and accumulated about 630 active minutes. Because the run counts were not equal, those totals are diagnostic rather than a fair final score.

Measure this instead:
Cost per accepted result = token spend + paid tools + reruns + human review time + repair work, divided by the number of outputs that actually pass.

Effort level is part of the cost decision. Anthropic's documentation says effort controls response thoroughness, tool calls, and token use. High is the default; lower settings can reduce spend, while max is intended for work where the extra capability justifies unconstrained token use. A fair comparison records the effort setting instead of silently letting one model explore for hours.

The Most Revealing Result Was a Test Mistake

Nate intended to compare Fable and Opus on Google Snake, but accidentally ran Opus twice. One run ignored the five-game limit, played many batches, took almost two hours, and cost about $18. The other followed the task more closely, took about 70 minutes, and cost about $10.

That mistake makes the result invalid as a model comparison and valuable as an operations lesson. Agent runs are nondeterministic. A single result can exaggerate a model's reliability, creativity, speed, or cost.

For consequential workflows, run at least three trials. Add a hard game, tool-call, time, or spend limit. Save the trace. Score each output against the same rubric. A model that occasionally produces a brilliant artifact but often wanders may be worse than a slightly less impressive model with a dependable performance floor.

Practical Model Routing

Work type Start with Why Control
Bug fixes with deterministic tests Opus 5 Strong verification behavior and favorable result in the clearest code test. Require tests, minimal diff, and reviewer approval.
Bounded research and analysis Opus 5 at medium or high effort Good professional work at lower token rates. Require source URLs, quoted evidence, and a claim audit.
Visual concepts, decks, and creative direction Fable 5 Faster and preferred on several visual deliverables in this suite. Set factual inputs and a strict output budget.
Long, ambiguous multi-stage projects Fable as manager, Opus as worker Preserves Fable for planning and judgment while moving execution to the cheaper model. Give each subtask a definition of done and capped budget.
Routine copy or formatting Sonnet or a cheaper model Both Opus and Fable can be unnecessary for low-risk work. Escalate only when the cheaper lane fails the rubric.
Computer use Whichever model passes repeated trials The accidental double-Opus run showed large variance. Cap steps, time, spend, and irreversible actions.

Copy-Ready Evaluation Brief

Goal
Complete [real workflow] using the supplied project and tools.

Controlled setup
- Harness:
- Model:
- Effort:
- Tools and permissions:
- Starting context:
- Time limit:
- Token or dollar budget:
- Number of trials: 3

Acceptance tests
1. [Objective functional test]
2. [Accuracy or source test]
3. [Visual or usability rubric]
4. [Security and permission check]
5. [Required files and handoff]

Stop conditions
- Stop when every acceptance test passes.
- Stop and report the blocker when the budget is reached.
- Do not widen permissions or publish without approval.

Record
- Input, output, and cached tokens
- Tool calls and subagents
- Wall-clock and active time
- Repairs and human review minutes
- Pass or fail for every acceptance test

Decision
Choose the lowest-cost model that passes repeatedly.
Escalate only the tasks that need more intelligence or judgment.

Video Chapters

00:00Opus 5 benchmarks versus Fable 5
01:00Testing without a harness
03:17Codebase bug hunts
05:19Why verification matters
06:51AIS Live announcement video
09:37Landing page build
12:15LinkedIn post and carousel
14:38Audience research report
16:58Video outline and slide deck
20:18Computer-use Snake game
22:39Structural simulator build
27:41Total cost and token breakdown
29:19Final thoughts

Bottom Line

Opus 5 looks like an excellent default for serious daily work, especially when the task has objective tests and benefits from persistent self-verification. Fable 5 still earns its premium on ambitious orchestration and several kinds of visual or creative work. In Nate's suite, neither model won every lane.

The bigger lesson is operational. Token pricing, benchmarks, and one-shot demos are inputs, not the decision. Control the harness, record effort, run repeated trials, cap the agent, verify the result, and route the work according to the evidence your own business produces.

Sources

Common questions

Is Claude Opus 5 better than Claude Fable 5?
Not universally. In Nate Herk's tests, Opus 5 produced the strongest verified bug-fix result and a more usable structural simulator. Fable 5 was generally faster and was preferred for several visual and creative deliverables. The right choice depends on the task, effort level, harness, and acceptance tests.
Is Opus 5 cheaper than Fable 5?
Its standard API token rates are half of Fable 5: $5 per million input tokens and $25 per million output tokens versus $10 and $50. A complete Opus task can still cost more if it uses more output tokens, tool calls, subagents, retries, or verification loops.
Were the models tested without an agent harness?
One Excalidraw diagram test was run in Claude chat without skills or a verification path. Most experiments ran inside Claude Code, so they compare the models within the same harness rather than isolated raw model behavior.
What was the most convincing Opus 5 result?
The second codebase bug hunt had the clearest objective evidence. Opus passed four of four checks and received a 93 out of 95 technical score in Nate's Codex review, while Fable passed two of four and scored 66 out of 95.
Why was the computer-use result not a valid Opus versus Fable test?
Nate accidentally ran Opus 5 twice. The two runs followed the same prompt but took very different paths, costs, and times. That makes the result useful evidence of nondeterminism, but not a Fable comparison.
How should a team compare Opus 5 and Fable 5?
Use the same harness, tools, context, effort policy, timeout, budget, and acceptance tests. Run each model several times, record repair work and reviewer time, and compare cost per accepted result rather than token price or the best-looking first output.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call