AI Workflow Design

Self-Improving AI Agents: Duncan Rogoff's Overnight Loop

Duncan Rogoff's self-improving agents video shows a simple but powerful loop: make one thing, evaluate it against examples, preserve the lesson, and use that lesson next time. His companion free written guide supplies the file structure, six setup prompts, and a three-night web-page experiment. This article organizes both so you can see what is working, what is still experimental, and what to control before an agent runs unattended.

The short answer

A "self-improving" agent here is a scheduled workflow improving its own instructions, not a model retraining itself. Duncan's agent makes a draft, a critic compares it with fixed benchmarks, and the agent records a small, sourced rule for the next run. The useful unit of progress is an auditable change to the process. The proof still needs a stable rubric and, ultimately, a person or audience outside the agent's own grading loop.

Watch Duncan's Walkthrough

Credit: The system, examples, and reported scores belong to Duncan Rogoff. Read his original guide for the full six copy-paste setup prompts. The safety and measurement notes below are editorial additions, not claims that Duncan demonstrated a validated commercial result.

The Five-Part Overnight Loop

Duncan adapts the experimental rhythm of Andrej Karpathy's autoresearch: make a bounded change, run a test, keep it if it wins, discard it if it loses. Autoresearch uses an objective training metric; a web page or Reel needs a subjective judge, examples of good work, and human taste. That difference is crucial.

PartWhat it doesFailure to watch
One artifactMake the same kind of output each run: a page, short video, or draft.Changing the deliverable makes comparisons meaningless.
A judgeScore it against a stable rubric and strong reference examples.A friendly critic can make weak work look better.
A short rulebookKeep lessons in PLAYBOOK.md, each with the run or human note that supports it.Unbounded rules turn into noise or contradictions.
One focused experimentChange one axis per night, such as clarity or motion, then compare.Five simultaneous changes hide the cause of a result.
A bounded scheduleRun for a fixed time and cost, then leave an artifact and morning report.Unattended work can keep spending or change the wrong files.

Build the Judge Before the Builder

The guide's critic uses roughly eight to ten categories, each scored 0-10, and three to five human-selected benchmarks. Before assigning numbers it compares the new artifact beside three examples and asks which is stronger. Duncan caps some craft scores at 5 when a benchmark is visibly better. He suggests calibrating 5 as competent, 7 as benchmark quality, and 9 as clearly better. This is a discipline against inflated self-grading, not an objective guarantee of quality.

Keep the rubric fixed during a comparison. If a category or weighting changes, re-score the old best under the new rubric before declaring victory. Duncan's guide is unusually candid about this: his first run looked weak to him despite a generous internal score. A separate reviewer, blind side-by-side user preference test, or real publishing metric is stronger evidence than an agent judging its own output.

The improvement memory is deliberately small: PLAYBOOK.md has a 60-line cap, one lesson per line, and a source for each change. Human notes in FEEDBACK.md take precedence over the night's focus rotation. A retro either keeps the experiment with evidence or records why it was dropped. A git commit and MORNING.md make the changes inspectable in daylight.

The Guide, Organized as Six Setup Stages

Duncan's original guide has the exact prompts. These are the jobs they perform, in order; the early stages are manual on purpose.

StageOutputCheck before continuing
1. Interview and define qualityAUDIENCE.md, RUBRIC.md, benchmarks/, initial PLAYBOOK.md, GUARDRAILS.md, FEEDBACK.md, and focus-axis files.Can you tell a poor example from an excellent one using the written rubric?
2. Build the run recordCLAUDE.md, runs/, ledger.jsonl, MORNING.md, and a retro prompt.Are hard limits immutable and every change reversible?
3. Create the criticA separate critic prompt writing critique.md and score.json after benchmark comparison.Does it cite visible evidence and score bad work harshly?
4. Run one night by handPlan, artifact, critique, limited fixes, retro, ledger entry, commit, and morning report.Did the entire loop finish inside time and cost caps without external action?
5. Compress the reportA three-line MORNING.md: score, change, and what you need to review.Can a person understand what happened without reading a long transcript?
6. Schedule only after two clean manual runsA local runner and scheduler; Duncan's example uses macOS launchd.Confirm paths, permissions, power state, budget, and a way to stop the job.

Claude Code's current non-interactive documentation explains claude -p, CLI options, and unattended permission behavior. Check those docs and settings precedence before copying the guide's runner: tool flags and defaults can change, and a scheduled process may load more project context and integrations than expected.

What the Three Web-Page Nights Actually Showed

RunDuncan's reported scoreHow to read it
Night 1, October 337 to 43 to 45 on an older 60-point rubricHe disliked the page despite the score. Do not compare this scale with later runs.
Hand-built benchmark72/90 after earlier iterationsA comparison page built by Duncan, not a public audience preference test.
Night 2, October 456 to 64/90 across eight roundsIt did not beat the 72-point bar; a losing night was recorded.
Night 3, October 573 to 77/90 across five scoresHis critic judged the final page above the 72-point benchmark. The run took about an hour and used five images.

Those numbers come from Duncan's guide. He later added a tenth category, making future scores out of 100. The observed improvement is within his own evaluation process; it does not yet show that visitors preferred the page, that customers converted, or that every later run improves. One of the three nights missed the bar.

Where He Applies the Loop, and What Is Still New

  • Web pages: the clearest worked example. The agent researches a topic, writes a brief, builds the page and assets, captures screenshots, runs critique/fix rounds, then updates its rulebook.
  • Instagram Reels: an early restyle lost to the original; a later staircase animation beat a simpler ladder-style visual in his internal comparison. The lesson was about making the idea clearer, not merely adding effects.
  • Substack writing: he is beginning to train the system on his stories, opinions, and feedback. The written guide explicitly says this loop is new and has no results yet.
  • Long-form video: discussed as a next application, not a proven overnight production line. Each format needs its own quality standards and review process.

What to Lock Down Before It Runs While You Sleep

  1. Keep authority narrow. Let the agent draft inside a dedicated folder. Do not grant publishing, sending, purchasing, deployment, or account-change authority to the overnight run. Treat morning review as a required handoff.
  2. Use real containment. Duncan notes that a broadly allowed shell can sidestep text-based edit deny rules. Prompt instructions and file deny-lists are not a security boundary. Use a disposable or isolated workspace, minimal credentials, restricted network/tools, and OS-level limits where appropriate.
  3. Cap time and spend outside the prompt. Put hard limits on runtime, model usage, and paid media generation. Stop on errors instead of retrying indefinitely.
  4. Preserve evidence. Keep benchmarks and hard limits human-owned. Log input, output, rubric version, score, changed rule, cost, and commit. Only compare runs under the same rubric or re-score the earlier one.
  5. Validate with people. Periodically blind-review samples, check the actual audience or business outcome, and reject rules that improve the critic's score while making the work worse.

The first pilot can be entirely manual. If it cannot produce a useful artifact, an honest critique, and a readable morning report twice in a row, scheduling it only automates the confusion.

Watch by Topic

TimeDiscussion
00:00What Duncan means by a self-improving agent
01:13Karpathy's autoresearch inspiration
04:01The overnight web-page workflow
05:28Training the critic's taste with examples
10:42Rules the agent writes for its next run
11:05Short-video experiments
13:01Human feedback file
14:03Substack writing loop

A Business Idea Built From the Loop

The sellable version is not "an AI team that publishes while you sleep." It is a quality-improvement service for one repeated deliverable, with an owner-approved rubric and a human sign-off. Use this prompt to test whether a buyer would pay for that narrower outcome.

Business idea prompt

Test an overnight quality-loop service

A seven-day manual pilot before any scheduled agent.

Ready to copy

Sources and Further Reading

Common questions

Does the agent retrain Claude Opus 5.5?
No. In this workflow the agent changes instructions, craft rules, and a local playbook between runs. It does not update the model weights.
Did the creator prove the agent improves audience outcomes?
Not yet. Duncan reports internal critic scores for three web-page nights, including a 77/90 result against a 72/90 hand-built benchmark. These are his judge scores, not independent user or business outcomes.
What should remain under human control?
The immutable task and audience definition, permission boundary, benchmark set, budget, and any publishing, messaging, spending, or deployment decision. Review artifacts and rule changes before using them externally.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call