AI Workflow Design

How to Run a Company With Claude Code: James McAulay's Stack

Direct Answer

James McAulay is not handing a company to one autonomous chatbot. He has made Claude Code the operating surface through which a founder can search company memory, create content, inspect analytics, build software, configure email, update a CRM, and dispatch narrower cloud agents. That distinction matters. The impressive part is not a model acting as CEO. It is an intentionally organized business becoming legible and operable through files, skills, APIs, tests, and approvals.

In Andrew Warner's interview, James demonstrates a complete LinkedIn lead-magnet loop: Claude researches the idea, asks questions, creates a wireframe, produces an animated infographic, builds a landing page, connects a Kit form and sequence, and uses PostHog data to learn what converted. The same environment can work with GitHub, Next.js, Vercel, Sentry, Cubic, Playwright, Attio, Apollo, Granola, and a large Markdown knowledge base.

The useful thesis: give an agent an owned knowledge base, a small library of reviewed skills, narrowly scoped connections, observable feedback, and a promotion path from draft to approved action. Claude becomes an operating layer. The founder still owns intent, money, risk, relationships, and the final decision.

Watch the Interview

Credits: the workflow, demonstrations, and business figures come from Andrew Warner's interview with James McAulay. James runs Agent Accelerator and published the LinkedIn Infographic skill shown in the episode. Tool behavior and deployment boundaries were checked against official product documentation on 5 September 2026.

What the $250K-a-Month, Zero-Employee Claim Means

The interview describes James's business as generating roughly $250,000 per month with zero employees. Treat both parts as founder-reported context, not audited evidence. The episode does not show financial statements, recurring versus one-time revenue, refunds, gross margin, contractor involvement, customer concentration, or the time James spends operating and improving the system.

"Zero employees" is also a narrow accounting description. It does not mean zero human labor. James supplies judgment, brand taste, customer knowledge, corrections, and approvals. The stack relies on product teams and infrastructure at Anthropic, Vercel, PostHog, Kit, Sentry, CRM vendors, model providers, and other services. Legal, tax, design, sales, or customer support may still involve outside humans even when nobody is on payroll.

ClaimWhat the video supportsWhat it does not establish
$250K per monthFounder-reported business scaleAudited revenue, margin, stability, or attribution to Claude
Zero employeesNo conventional employee team is describedZero contractors, vendors, founder labor, or professional support
Claude runs the companyClaude coordinates many recurring knowledge-work workflowsIndependent strategy, legal accountability, banking control, or unsupervised governance
One connected workspaceContent, code, data, analytics, email, and CRM can be reached from one interfaceThat every system should receive broad or permanent write access

The stronger evidence is visible in the workflow itself. Claude can move from an idea to an artifact, connect that artifact to distribution and capture, inspect the resulting data, and preserve the process as a reusable skill. That is a real operating advantage even without the headline number.

The Six-Layer Business Operating Stack

The demo looks fluid because the underlying responsibilities are separated. A useful reconstruction has six layers:

  1. Knowledge: a searchable LLM wiki of offers, audience insight, content, meetings, decisions, and operating history.
  2. Instructions: CLAUDE.md files, checklists, and reusable skills that define how a job should be done and what "done" means.
  3. Tools: selected APIs, MCP servers, browser access, code repositories, analytics, email, and CRM connections.
  4. Delivery: LinkedIn, landing pages, Kit forms, Resend messages, GitHub branches, and Vercel previews.
  5. Observation: PostHog events, session evidence, Sentry errors, Playwright tests, Cubic review, and business outcomes.
  6. Runtime: a local Claude Code session for attended work and a constrained cloud agent for schedules or remote channels.

This is closer to a company control plane than a universal employee. The model translates intent across systems, but each external system remains the system of record. Kit owns subscriber state, GitHub owns code history, Vercel owns deployment state, PostHog owns product events, and the CRM owns commercial records. Claude should never become the only place where a decision or action exists.

Design rule: keep the intelligence flexible and the records deterministic. Let the model propose copy, queries, layouts, classifications, and plans. Let APIs, schemas, tests, permissions, and humans decide what is accepted and committed.

The LinkedIn Infographic Is a Full Revenue Loop

The episode's clearest example starts with an infographic, but the real product is the loop around it. James uses voice input through Wispr Flow to explain an idea quickly. Claude then uses his published infographic skill to convert that intent into a sequence of reviewable artifacts.

  1. Brief: define the audience, desired lesson, evidence, CTA, brand constraints, and destination.
  2. Clarify: ask questions where a missing choice would materially change the output.
  3. Wireframe: show information hierarchy before spending time on final styling or animation.
  4. Approve: a human checks the thesis, claims, examples, sequence, and visual direction.
  5. Produce: generate the infographic in HTML and CSS, add motion, and export a GIF for LinkedIn.
  6. Convert: build a matching landing page, form, tag, lead magnet, and Kit email sequence.
  7. Measure: record post engagement, page visits, form completion, activation, and downstream revenue.
  8. Learn: update the skill only from reviewed outcomes, not vanity engagement alone.

The wireframe step is doing more work than it appears. Low-fidelity review makes structural corrections cheap. A reviewer can change the order, remove a weak claim, or sharpen the CTA before animation makes the artifact feel finished. The same logic applies to landing pages and email sequences: approve the message architecture before polishing production output.

Animated GIFs also need a publishing contract. Verify dimensions, file size, first-frame clarity, contrast, text size, playback speed, reduced-motion alternative, link destination, and a static fallback. LinkedIn engagement is not the final KPI. The post should carry a tagged URL so the business can distinguish attention from subscribers, qualified conversations, and sales.

The LLM Wiki Is the Compounding Asset

James describes an LLM wiki containing roughly 2,000 to 3,000 Markdown files and about eight months of operating history. That repository is what lets Claude answer with company-specific context instead of generic marketing language. It can retrieve the offer, audience vocabulary, prior experiments, analytics notes, brand patterns, and decisions behind the current task.

Volume alone is not a knowledge system. Thousands of files can create confident contradiction if the model finds an old offer beside a current one, a brainstorm beside a decision, or a customer quote without consent or provenance. Each durable document should expose enough metadata to be governed:

  • Owner, status, created date, last reviewed date, and next review date.
  • Source links and whether the statement is fact, hypothesis, decision, or draft.
  • Scope such as company, product, campaign, customer, or personal.
  • Sensitivity such as public, internal, confidential, or restricted.
  • Supersedes and superseded-by references for changed decisions.
  • Retention and deletion rules for customer, employee, and personal data.

Keep credentials outside the wiki. API keys, session cookies, private tokens, recovery codes, production secrets, raw payment data, and unnecessary personal information belong in a secret manager or the source system, not in model-readable Markdown. Add index files so an agent can find the current operating truth before searching historical material.

The most valuable documents are often not polished essays. Decision logs, failed-test notes, accepted examples, customer-language snippets with permission, metric definitions, and postmortems teach an agent how the company thinks. This is how the workspace compounds instead of restarting from one giant chat.

PostHog Turns Output Into a Feedback Loop

Connecting PostHog lets Claude inspect product and funnel data, ask what is performing, and propose experiments. That is more useful than asking a model to "improve conversion" from a screenshot because the agent can work from defined events, cohorts, paths, and page behavior.

The model still needs a measurement contract. Before it touches a page, define the primary conversion, eligibility window, guardrail metrics, minimum sample, attribution rule, and decision owner. A rise in form submissions may be meaningless if lead quality falls, unsubscribes increase, or tracking changed during the test.

LayerUseful metricCommon trap
LinkedIn postQualified profile visits and tagged clicksOptimizing for impressions or comments alone
Landing pageEligible visitor-to-confirmed-subscriber rateCounting bots, staff, repeat visits, or unconfirmed forms
Email sequenceDelivery, replies, activation, and unsubscribe rateTreating opens as reliable intent
ProductActivation and retained useShipping UI changes without behavioral evidence
RevenueQualified pipeline and collected contribution marginAttributing every later sale to the first visible touch

Claude can summarize evidence and draft the next test. It should not silently rewrite metric definitions, compare mismatched periods, or ship a variant because one noisy chart moved. Save the query, date range, segment, and source event beside every recommendation.

Build, Review, Preview, Then Deploy

The software loop connects a Next.js codebase to GitHub and Vercel. Claude can implement a landing page, run it locally, inspect the flow, and prepare a deployment. The mature version uses several independent checks rather than asking the same model that wrote the code whether its work is good.

  1. Plan: identify files, behavior, data flow, acceptance criteria, risks, and rollback.
  2. Implement: work on a branch with a small diff and no unrelated changes.
  3. Test: run type checks, linting, unit tests, and targeted integration tests.
  4. See: use Playwright to exercise the actual desktop and mobile journey, including console and network failures.
  5. Review: use the diff plus a separate review layer such as Cubic, then resolve findings with human judgment.
  6. Preview: deploy to an isolated Vercel preview with synthetic or safe test data.
  7. Approve: a human verifies copy, visual quality, accessibility, analytics, privacy, and the critical path.
  8. Promote: release to production, watch Sentry and product telemetry, and preserve a rollback target.

AI review is another signal, not an approval authority. The official Cubic workflow can comment on pull requests and propose fixes; its documentation still expects a developer to inspect and accept the result. Sentry finds runtime failures after code exists. Playwright checks flows you explicitly test. None of them replaces architecture review, secrets scanning, dependency policy, backups, or a person accountable for production.

Email and CRM Need Separate Write Boundaries

The lead-magnet workflow uses Kit for forms, tags, subscribers, and sequences. Kit's current API supports list management, tags, custom fields, broadcasts, and related automation. The demo also separates transactional messages through Resend, which is a sensible operational boundary: newsletter behavior should not share every sending path, credential, or reputation dependency with account-critical mail.

For CRM work, James connects Attio and Apollo so new leads can be enriched and prioritized. That can save time, but it creates data-governance work. The system should retain only fields needed for a defined purpose, preserve source and timestamp, respect provider licenses and applicable privacy law, handle suppression and deletion, and avoid turning probabilistic enrichment into a fact about a person.

Safe promotion path: read records, draft a proposed update, create a review queue, allow one reversible field update, then expand only after accuracy and downstream effects are measured. Keep bulk sends, deletions, ownership changes, billing, and sequence activation human-owned.

Every write should have an idempotency key or duplicate check, a before-and-after record, an initiating user or schedule, and a recovery path. The agent should stop when identities conflict, consent is unclear, the source is stale, or a requested action would cross from marketing into a regulated or contractual decision.

Vercel Eve Moves the Agent From Laptop to Runtime

James demonstrates a cloud agent named Jamie that delivers a daily sales and pipeline briefing through Telegram. The enabling idea is Vercel Eve, an open-source, filesystem-first framework for durable AI agents. An Eve project can define instructions, tools, skills, channels, schedules, evaluations, and human-in-the-loop behavior in version-controlled files.

That makes an existing agent folder portable into a runtime, but not production-ready by declaration. Eve is currently marked beta, and its deployment guide explicitly requires replacing placeholder authentication with a production route policy. It also expects proper model credentials, database and object storage, channel secrets, and verification of the deployed health and agent routes.

  • Give the cloud agent a separate service identity, not the founder's universal credentials.
  • Allowlist tools and destinations per job; deny everything else by default.
  • Store secrets in managed environment settings and rotate them.
  • Set model, token, time, concurrency, and monetary budgets.
  • Require approval for external messages and consequential writes.
  • Log tool calls, source records, outputs, approvals, failures, and retries.
  • Run evaluations against representative cases before every instruction or model change.
  • Test pause, revocation, duplicate prevention, and rollback before enabling schedules.

A daily briefing is a good first cloud job because the output is read-only and the recipient is known. A daily agent that edits CRM stages, launches campaigns, or changes production code carries a very different risk class and should not inherit the same permission profile.

The Health-Agent Example Needs a Harder Privacy Boundary

The interview extends the architecture to personal fitness data and a health-oriented agent. The technical pattern is understandable: synchronize user-owned sources, turn large exports into structured tables or summaries, and ask the agent to surface trends. The sensitivity is much higher than a content calendar.

Health-adjacent data can reveal identity, location, sleep, activity, reproductive information, routines, conditions, and other intimate patterns. A consumer fitness export is not automatically protected by the same rules as a hospital record, and every connector can create another copy. Use data minimization, explicit consent, encryption, short retention, provider review, access logs, export and deletion controls, and a separate workspace from company operations.

The agent should summarize and help the owner prepare questions. It should not diagnose, prescribe, change treatment, contact a clinician, or infer high-consequence conditions without a qualified professional. Do not route this data through a general business knowledge base merely because the same framework can read it.

A Permission Model for a Founder-Led AI Company

ClassExamplesDefault
ObserveRead docs, analytics, public pages, approved CRM viewsAllow with scope, logging, and retention limits
DraftCopy, wireframes, code branches, email drafts, proposed CRM changesAllow in isolated destinations
Reversible writeCreate preview, add internal note, apply a tested tagApprove initially; automate only after measured reliability
External actionPublish post, send email, message lead, deploy productionHuman approval with exact preview
Irreversible or regulatedDelete records, move money, sign terms, change medical or legal stateHuman-owned; agent may prepare evidence only

Skills deserve the same scrutiny as code. Agent Accelerator's own curriculum warns that an installed skill is code the user is about to run. Review instructions and supporting scripts, pin versions, limit dependencies, test in a sandbox, and record who approved the release. A useful skill repository has owners, changelogs, tests, deprecation rules, and rollback, not just a folder of prompts.

The operating dashboard should track cost per accepted artifact, human review time, correction rate, stale-source rate, failed tool calls, unauthorized-action attempts, duplicate writes, production incidents, subscriber quality, unsubscribe rate, and collected contribution margin. The goal is not maximum autonomous activity. It is more accepted work per unit of founder attention without increasing risk or customer confusion.

A 30-Day Rollout

WeekBuildControlPass condition
1Create a small, current knowledge base and one content skillNo external writes; remove secrets and sensitive dataTen historical tasks produce useful, source-traceable drafts
2Connect read-only analytics and create a local preview flowMetric contract, branch isolation, deterministic testsRecommendations reproduce from saved queries and previews pass review
3Add one Kit sandbox or CRM review queueDraft-only actions, duplicate checks, before-and-after logTwenty cases meet accuracy and consent requirements with no silent writes
4Deploy one read-only Eve briefing on a scheduleService identity, budgets, allowlist, evaluations, kill switchSeven consecutive runs arrive on time, cite sources, and fail safely

Start with the content loop because the artifacts are visible and reversible. Do not start by granting a cloud agent broad access to production, customer records, email sending, and personal data at once. Every successful layer should earn the next permission through evidence.

A compact agent contract should name the job, sources, expected output, acceptance test, allowed tools, destination, write mode, budget, timeout, escalation triggers, retention, rollback, and accountable owner. That one page turns "Claude runs the company" into something a real company can inspect.

Video Chapters

TimeTopicTimeTopic
00:00Claude Code as the business interface12:18Lead-magnet checklist
00:42LinkedIn infographics12:37Kit forms, tags, and sequences
02:15The reusable infographic skill14:28Animated GIF export
02:32Voice input with Wispr Flow15:44Local landing-page preview
03:14Lead-magnet funnel16:25Kit API connection
03:37PostHog performance data17:18Transactional email with Resend
04:26Clarifying questions17:50Attio and Apollo enrichment
05:34Wireframes before production19:03Daily Telegram briefing
06:40The LLM wiki20:03Vercel Eve cloud agents
07:53SEO and Lighthouse20:41Agent channels
08:44Next.js and GitHub21:23Organizing agent context
10:02Vercel deployment22:01Health-agent example
10:25Sentry and Cubic review25:37Structured data
10:46Playwright verification26:24Always-on cloud workflow
11:09Conversion optimization

Verdict

The episode is persuasive because it shows a connected operating loop, not a collection of isolated prompts. Company memory informs a reusable skill. The skill produces a campaign asset. The asset connects to a landing page and email system. Analytics return evidence. Code and infrastructure can then be changed through a reviewed delivery path. A cloud runtime handles selected recurring work.

The phrase "Claude runs the company" is still too strong. James runs the company through Claude. His judgment appears in the knowledge architecture, questions, wireframe approval, tool choices, metric interpretation, deployment checks, and permission boundaries. That is not a weakness in the model. It is the design that makes the model commercially useful.

For a founder, the first goal should not be a universal autonomous employee. Build one complete loop that remembers the business, produces an inspectable artifact, measures the outcome, and asks before acting outside its boundary. When that loop earns trust, add the next one.

Sources and Links

Editorial note: the business figures are attributed to James McAulay and have not been independently audited. Product capabilities, APIs, pricing, and beta status can change. This article was last checked on 5 September 2026 and is educational content, not legal, privacy, medical, financial, or security advice.

Common questions

Can Claude Code really run an entire company?
It can coordinate research, content, code, analytics, email setup, CRM enrichment, and scheduled workflows when those systems expose files, APIs, or browser interfaces. It does not own strategy, legal accountability, customer relationships, financial controls, or final quality. The founder remains the operator and accountable decision-maker.
Is the $250,000-a-month, zero-employee claim verified?
The figure and zero-employee description are reported in Andrew Warner's interview with James McAulay. The episode does not provide audited financial statements or a complete labor map. Zero employees does not mean zero humans: the business still depends on its founder, customers, software vendors, infrastructure providers, and potentially external professional support.
What is an LLM wiki?
In the demonstration it is a large, file-based business knowledge base containing thousands of Markdown documents. Claude can search it for offers, audience knowledge, past decisions, content, and operating context. A production version needs owners, dates, source links, sensitivity labels, retention rules, and a process for resolving conflicting or stale documents.
How does the LinkedIn infographic skill work?
The workflow starts with a topic and source context, asks clarifying questions, proposes a wireframe, waits for review, creates the visual in HTML and CSS, adds animation, and exports a GIF. It can then produce a linked landing page and Kit form or sequence. The reusable skill preserves the process, but the human still approves claims, design, audience fit, and publication.
Can Vercel Eve keep an agent running in the cloud?
Eve is a filesystem-first framework for durable agents with tools, skills, channels, schedules, and human-in-the-loop steps. It can be deployed to Vercel or a Node environment, but it is currently beta. Production use still requires real authentication, secret management, constrained tools, logs, evaluations, spending limits, and tested rollback.
Should an AI agent have direct access to Kit, Attio, Apollo, or production deployment?
Begin with read-only access and draft outputs. Promote one narrowly defined write at a time, such as creating a draft tag or opening a preview deployment. Keep bulk email, record deletion, billing, DNS, production promotion, and irreversible CRM changes behind explicit human approval and an audit log.
Is it safe to build a personal health agent from fitness data?
Treat fitness and health-adjacent data as sensitive. Minimize collection, separate identifiers, encrypt storage and transport, limit retention, review every connector, and keep medical interpretation with qualified professionals. A productivity agent should summarize user-owned data, not diagnose, prescribe, or silently send it to unrelated services.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call