Direct Answer
Claude can now coordinate most of an editorial motion-explainer pipeline from one detailed instruction. In Zubair Trabzada's demonstration, Claude Opus 5 researches a topic, writes the narration, plans six visual blocks, calls Higgsfield through MCP for images, clips, voice, and optional music, then assembles a finished short-form video.
That is meaningful automation, but it is not a hands-free publishing system. The model can produce a coherent first cut. A responsible creator still owns the thesis, evidence, source quality, visual rights, voice consent, pacing, captions, brand distinction, final edit, and decision to publish.
Watch the Workflow
Video credit: Zubair Trabzada | AI Workshop. The free prompt pack and skill file described in the video are distributed through Zubair's AI Workshop Light community linked from the YouTube description. This article does not reproduce that pack. The reusable prompt below is an original JQ AI SYSTEMS production framework.
Source and Terminology Note
The supplied transcript is the source for the demonstrated workflow, formats, example topics, voice-selection sequence, and Zubair's assessment of the results. Product capabilities were checked on 28 July 2026 against Higgsfield's official Skills, pricing, and terms pages; Anthropic's Claude Code Skills and MCP documentation; ElevenLabs' voice-cloning documentation; and YouTube's monetization and AI-disclosure policies.
"Vox-style" is used here as audience shorthand for an evidence-led editorial explainer with paper collage, bold typography, annotated maps, charts, visual metaphors, and documentary pacing. Vox did not sponsor or approve this workflow. A production system should encode visual principles, not clone another publisher's protected branding or assets.
The video's "100% automated" title describes the generation run, not the full publishing responsibility. Research tools can return weak sources. Generated charts can misrepresent data. Voice and music can introduce rights issues. Automated pacing can still feel repetitive. Those are not edge cases; they are the work of editorial production.
Useful Links
| Resource | Why it matters | Watch for |
|---|---|---|
| Zubair's full walkthrough | The primary demonstration, setup sequence, format changes, voice options, and finished examples. | The title is a creator framing, not a guarantee that review and editing disappear. |
| AI Workshop | Zubair's channel and the route to the free community resources mentioned in the episode. | The Jarvis assistant is described as a separate paid-community resource. |
| Higgsfield Skills | Official MCP capabilities, supported agents, model catalog, generation history, and authentication. | Higgsfield says MCP uses account credits and supports asynchronous generation. |
| Higgsfield pricing | Current plans and credit allowances. | Check it immediately before budgeting because plans, included credits, and model costs can change. |
| Higgsfield terms | Commercial use, output responsibility, prohibited uses, provenance, and subscription renewal. | Higgsfield does not warrant that outputs are original, legal, accurate, or non-infringing. |
| Claude Code MCP guide | How Claude Code connects to remote and local tool servers and manages authentication. | Review the server, scope, permissions, and status before granting access. |
| Claude Code Skills guide | How to turn a repeated procedure into a SKILL.md with supporting files and validation. | Keep the reusable method separate from client facts and one-off topics. |
| ElevenLabs voice-cloning documentation | Instant versus professional cloning, sample quality, verification, and limitations. | Use your own verified voice or an authorized shared voice. |
| YouTube monetization policy | Current rules for original, authentic, reused, repetitive, and mass-produced content. | Generic template output without original value can be ineligible for monetization. |
| YouTube AI disclosure guide | When creators must disclose realistic synthetic or meaningfully altered media. | Animated explainers may not always require a label, but realistic depictions of events can. |
| W3C captions guidance | Why narration, meaningful sound, and speaker information need synchronized text alternatives. | Burned-in decorative text is not a substitute for accurate captions. |
Episode Guide
| Time | Chapter | Practical takeaway |
|---|---|---|
| 00:00 | One-prompt examples | Three editorial styles show the target: a current-affairs explainer, a sports story, and an AI-business story. |
| 01:15 | Tools | Claude Code is the orchestrator; Higgsfield supplies generated media and audio capabilities. |
| 01:51 | Connect MCP | Add the official server, authenticate through the account flow, and verify the connector is enabled. |
| 03:48 | Prompt pack | The creator distributes several niche and format templates through a free community. |
| 05:18 | Run the 60-second prompt | Research precedes narration and media generation; the prompt uses six blocks of roughly ten seconds. |
| 09:11 | Skill shortcut | A SKILL.md stores the production procedure so later requests can be much shorter. |
| 12:08 | Formats and cost | The workflow can switch between vertical and landscape, but longer videos and retries consume more credits. |
| 14:00 | Voice and cloning | Choose a stock voice or use a properly authorized custom voice. |
| 15:54 | Results | The finished examples combine narration, zooms, collage, data points, transitions, and optional music. |
What the Workflow Actually Automates
| Production job | Agent contribution | Human ownership |
|---|---|---|
| Topic research | Search, summarize, compare sources, and extract candidate facts. | Choose credible sources, verify every consequential number, and reject misleading framing. |
| Editorial angle | Propose hooks, tension, chronology, and a closing thought. | Select the thesis and decide what the audience should understand. |
| Script | Draft narration to a target duration and reading pace. | Edit for accuracy, clarity, originality, tone, and fair representation. |
| Storyboard | Split the script into visual beats and specify charts, maps, collage, typography, and transitions. | Art-direct the visual logic and remove generic or derivative choices. |
| Asset generation | Call image, video, voice, and music tools through MCP. | Approve references, identities, rights, model choice, spend, and output quality. |
| Assembly | Sequence media, synchronize narration, add titles, and render a draft. | Check timing, readability, captions, mix, factual alignment, and export quality. |
| Publishing | Prepare files, titles, descriptions, chapters, and disclosure notes. | Make the release decision and accept responsibility for the published work. |
The right mental model is a producer with a tool belt, not a magic text-to-video box. Claude holds the plan and delegates media tasks. Higgsfield returns assets. The skill stores the method. Your review loop determines whether the result is useful.
The Six-Stage Production System
- Define the editorial contract. Set the audience, one-sentence thesis, evidence standard, runtime, format, tone, and forbidden shortcuts before research begins.
- Build an evidence sheet. Require a table containing each factual claim, source URL, publication date, confidence, and exact narration line that uses it.
- Lock narration before expensive media. Read the script aloud, cut weak lines, confirm timing, and approve the story before image or video credits are spent.
- Create a shot manifest. For every scene, define duration, purpose, on-screen text, visual source, motion, transition, audio cue, and rights status.
- Generate in passes. Make low-cost stills or previews first. Approve composition and continuity before requesting higher-resolution clips, voice, or music.
- Verify the rendered file. Watch once for story, once muted for visual comprehension, once audio-only for narration, and once at mobile size for text legibility.
Original Production Prompt
This is an original framework for an editorial collage explainer. It deliberately adds evidence, budget, rights, accessibility, and approval controls that a viral one-prompt demo can omit.
You are the editorial producer, motion director, researcher, and video engineer.
OUTCOME
Create a [60-second / 90-second / 3-minute] editorial motion explainer about:
[TOPIC]
AUDIENCE
[WHO IT IS FOR]
THESIS
The viewer should leave understanding:
[ONE SENTENCE]
FORMAT
- Canvas: [9:16 vertical / 16:9 landscape]
- Resolution: [1080x1920 / 1920x1080]
- Runtime tolerance: +/- 2 seconds
- Narration pace: 135-155 words per minute
- Captions: required, sentence case, safe margins
EVIDENCE
1. Research before scripting.
2. Prefer primary sources and current official data.
3. Create evidence.csv with:
claim, source, publication_date, confidence, narration_line.
4. Do not include a consequential claim with fewer than two credible sources
unless it comes from the authoritative primary source.
5. Mark disputed or estimated figures in the narration.
STORY
Use this arc:
1. Concrete hook
2. Context
3. Mechanism
4. Evidence
5. Consequence
6. Closing insight
VISUAL LANGUAGE
- Editorial paper collage
- Bold but restrained typography
- Annotated maps and diagrams
- Charts that begin at honest baselines
- Texture used for hierarchy, not noise
- No copied publisher logos, templates, footage, or branded assets
- Keep every scene visually distinct but inside one color and type system
PRODUCTION
1. Draft narration and storyboard only.
2. Stop for approval before paid media generation.
3. After approval, create shot-manifest.csv.
4. Use Higgsfield only for approved shots.
5. Generate previews before high-cost final assets.
6. Keep every source asset and generation record in /provenance.
7. Assemble, caption, mix, and render the draft.
AUDIO
- Use [approved stock voice / my verified voice: NAME].
- Never clone or imitate another person.
- Keep music below narration and avoid unlicensed reference tracks.
- Include a no-music export.
BUDGET
- Maximum image generations: [N]
- Maximum video generations: [N]
- Maximum retries per shot: [N]
- Maximum Higgsfield credits: [N]
- Ask before exceeding any limit.
VERIFICATION
Before declaring complete:
- Check every narration claim against evidence.csv.
- Check charts against their source values.
- Check names, dates, units, and pronunciation.
- Check captions and safe margins.
- Check visual continuity and repeated assets.
- Check that no third-party brand or protected asset is imitated.
- Export a review report listing unresolved risks.
DELIVERABLES
- final.mp4
- final-no-music.mp4
- captions.srt
- narration.txt
- evidence.csv
- shot-manifest.csv
- provenance/README.md
- review-report.md
The stop after scripting is intentional. Without it, the agent can spend credits producing polished visuals for a weak or inaccurate story.
Turn the Procedure Into a Claude Skill
Anthropic's official guide recommends a skill when you repeatedly paste the same instructions or maintain a multi-step procedure. The smallest useful structure is:
.claude/skills/editorial-motion-explainer/
|-- SKILL.md
|-- templates/
| |-- evidence.csv
| |-- shot-manifest.csv
| `-- review-report.md
|-- references/
| |-- visual-system.md
| |-- narration-style.md
| `-- rights-policy.md
`-- scripts/
|-- validate-evidence.js
|-- check-duration.js
`-- verify-deliverables.js
Keep the skill procedural. Do not hardcode one news story, one client's secrets, or one voice identifier into it. Store those in the project brief. A reusable skill should explain how to work, what to verify, what tools may be called, when to stop for approval, and what files define completion.
Before installing a community skill, read the complete SKILL.md and every referenced script. Look for hidden network calls, destructive commands, broad file access, credential handling, unexpected upload behavior, and pre-approved tools. A convenient skill is executable process knowledge, so treat it like code.
9:16 and 16:9 Need Different Direction
| Decision | 9:16 short-form | 16:9 long-form |
|---|---|---|
| Composition | One dominant subject; vertically stacked evidence; aggressive cropping. | Layered maps, charts, side-by-side comparisons, and wider environmental scenes. |
| Text | Five to nine words per card; larger type; central safe zone. | Short labels plus axes, legends, and supporting annotations. |
| Pacing | A meaningful visual change every one to three seconds. | Longer holds when the viewer needs to read or compare evidence. |
| Narration | One idea, minimal setup, one memorable conclusion. | More context, counterargument, methodology, and source nuance. |
| Captions | Keep clear of platform UI and lower-third controls. | Leave space for player controls and avoid covering charts. |
| Reuse | Design natively; do not simply crop the landscape master. | Build a master evidence package that can feed shorter derivatives. |
The transcript shows the agent correcting a preset that initially intercepted the requested landscape output. That is a useful reminder: state width, height, aspect ratio, caption safe area, and export resolution explicitly. "Make it horizontal" is not a production specification.
Cost Controls That Prevent Credit Surprises
Higgsfield says MCP generations consume the same account credits as the platform and that the cost varies by model and resolution. The transcript does not provide a stable per-video price, so a responsible estimate should be built from the actual shot manifest.
estimated_run_cost =
research_and_model_usage
+ (image_previews x preview_cost)
+ (final_images x final_image_cost)
+ (video_clips x clip_cost)
+ voice_generation
+ music_generation
+ retries
+ rendering_and_storage
Track cost per approved minute, not cost per generation. A cheap clip that fails the brief is waste. A more expensive asset that survives review and can be reused may be the better buy.
- Set a hard credit ceiling before the run.
- Require approval after research, script, storyboard, and preview passes.
- Cap retries per scene and log the reason for each retry.
- Reuse approved characters, textures, maps, and visual systems from generation history.
- Render a ten-second style test before commissioning a full minute.
- Keep a no-music version so a weak soundtrack does not force a full rebuild.
Voice and Cloning
Zubair demonstrates a stock Higgsfield voice and describes creating a custom voice through an ElevenLabs option. ElevenLabs distinguishes quick instant clones from higher-consistency professional clones and uses verification as an ethical and legal safeguard.
The operational rule is simple: clone your own verified voice or use a voice explicitly shared by its owner through the provider's supported process. Do not upload interviews, podcasts, employee recordings, celebrity clips, or client calls to imitate a person. A technically possible clone is not automatically an authorized one.
For production, generate a pronunciation sheet for names, acronyms, places, currencies, and technical terms. Listen for emotional mismatch, unnatural emphasis, hallucinated words, and pacing that leaves too little time for charts. Keep the narration script and voice settings with the project so revisions remain reproducible.
Seven Human Review Gates
- Thesis gate: is the central claim useful, fair, and specific?
- Evidence gate: does every important number match a current source, unit, date, and denominator?
- Rights gate: are footage, images, fonts, music, voices, logos, and references cleared for the intended use?
- Design gate: does the piece have its own visual system rather than a generic template or a close imitation of another publisher?
- Accessibility gate: are captions accurate, synchronized, readable, and separate from decorative on-screen text?
- Technical gate: are resolution, frame rate, audio peaks, safe margins, codec, file size, and platform exports correct?
- Release gate: has a named human watched the complete final export and approved the title, thumbnail, description, sources, and AI disclosure?
For news or geopolitics, add an eighth gate: an editor should compare the finished visual sequence with the source record. A generated image can make an unverified event look documented. That is a materially different risk from a typo in narration.
Faceless Channels Still Need a Face
The face can be a point of view rather than an on-camera person. YouTube's current monetization policy says repetitive, mass-produced, or templated AI content without meaningful original value can be ineligible. It explicitly rewards original commentary, narrative, educational value, and visible creative participation.
A durable channel therefore needs an editorial signature:
- Original research or a clearly argued interpretation.
- A consistent but evolving visual language.
- Sources in the description and corrections when facts change.
- Human-written observations that are not interchangeable with another channel.
- Licensed media and music records.
- AI disclosure when realistic synthetic media could mislead viewers about a person, place, or event.
Automation should make originality more affordable, not make sameness faster.
A Seven-Day Pilot
| Day | Build | Acceptance test |
|---|---|---|
| 1 | Choose one evergreen topic and collect five primary sources. | A human can explain the thesis and evidence without the model. |
| 2 | Write and time a 45- to 60-second script. | The narration fits the runtime and every claim appears in the evidence sheet. |
| 3 | Create six scene cards and one visual-system page. | Every visual teaches something; none exists only to fill time. |
| 4 | Connect Higgsfield MCP and produce one ten-second style test. | The connector, save location, credit tracking, and visual direction all work. |
| 5 | Generate approved assets and assemble the full draft. | The run stays inside the credit and retry budget. |
| 6 | Review facts, rights, voice, captions, mix, pacing, and mobile readability. | The review report contains no unresolved high-risk item. |
| 7 | Publish privately or unlisted, collect five viewer reactions, and revise the skill. | The next run incorporates concrete lessons rather than repeating the same template. |
Bottom Line
Claude Opus 5, Claude Code, and Higgsfield MCP can compress a multi-tool motion workflow into one coordinated agent run. That is a real production advantage. The reusable skill is even more valuable because it turns one successful experiment into a repeatable system.
But the most important file is not the final video. It is the review standard around it. When the system records its evidence, shot logic, provenance, budget, captions, and unresolved risks, automation becomes dependable. Without those controls, it merely produces polished uncertainty at speed.
Sources
- I 100% Automated Vox-Style Motion Graphics With Claude - Zubair Trabzada's primary walkthrough.
- Higgsfield Skills - official MCP capabilities, supported clients, model access, generation history, authentication, and credit behavior.
- Higgsfield pricing - current plans and credit allowances.
- Higgsfield Terms of Use - commercial use, output responsibility, provenance, restrictions, and subscription terms.
- Extend Claude with skills - official
SKILL.mdstructure, locations, supporting files, and invocation controls. - Connect Claude Code to tools via MCP - official server setup, authentication, status, and scope guidance.
- Voice cloning: how it works - official cloning methods, recording quality, verification, and limitations.
- YouTube channel monetization policies - original, repetitive, reused, and mass-produced content guidance.
- Disclosing use of generative AI content - official disclosure requirements and examples.
- W3C captions guidance - synchronized alternatives for narration and meaningful sound.