Duncan Rogoff's self-improving agents video shows a simple but powerful loop: make one thing, evaluate it against examples, preserve the lesson, and use that lesson next time. His companion free written guide supplies the file structure, six setup prompts, and a three-night web-page experiment. This article organizes both so you can see what is working, what is still experimental, and what to control before an agent runs unattended.
A "self-improving" agent here is a scheduled workflow improving its own instructions, not a model retraining itself. Duncan's agent makes a draft, a critic compares it with fixed benchmarks, and the agent records a small, sourced rule for the next run. The useful unit of progress is an auditable change to the process. The proof still needs a stable rubric and, ultimately, a person or audience outside the agent's own grading loop.
Watch Duncan's Walkthrough
Credit: The system, examples, and reported scores belong to Duncan Rogoff. Read his original guide for the full six copy-paste setup prompts. The safety and measurement notes below are editorial additions, not claims that Duncan demonstrated a validated commercial result.
The Five-Part Overnight Loop
Duncan adapts the experimental rhythm of Andrej Karpathy's autoresearch: make a bounded change, run a test, keep it if it wins, discard it if it loses. Autoresearch uses an objective training metric; a web page or Reel needs a subjective judge, examples of good work, and human taste. That difference is crucial.
| Part | What it does | Failure to watch |
|---|---|---|
| One artifact | Make the same kind of output each run: a page, short video, or draft. | Changing the deliverable makes comparisons meaningless. |
| A judge | Score it against a stable rubric and strong reference examples. | A friendly critic can make weak work look better. |
| A short rulebook | Keep lessons in PLAYBOOK.md, each with the run or human note that supports it. | Unbounded rules turn into noise or contradictions. |
| One focused experiment | Change one axis per night, such as clarity or motion, then compare. | Five simultaneous changes hide the cause of a result. |
| A bounded schedule | Run for a fixed time and cost, then leave an artifact and morning report. | Unattended work can keep spending or change the wrong files. |
Build the Judge Before the Builder
The guide's critic uses roughly eight to ten categories, each scored 0-10, and three to five human-selected benchmarks. Before assigning numbers it compares the new artifact beside three examples and asks which is stronger. Duncan caps some craft scores at 5 when a benchmark is visibly better. He suggests calibrating 5 as competent, 7 as benchmark quality, and 9 as clearly better. This is a discipline against inflated self-grading, not an objective guarantee of quality.
Keep the rubric fixed during a comparison. If a category or weighting changes, re-score the old best under the new rubric before declaring victory. Duncan's guide is unusually candid about this: his first run looked weak to him despite a generous internal score. A separate reviewer, blind side-by-side user preference test, or real publishing metric is stronger evidence than an agent judging its own output.
The improvement memory is deliberately small: PLAYBOOK.md has a 60-line cap, one lesson per line, and a source for each change. Human notes in FEEDBACK.md take precedence over the night's focus rotation. A retro either keeps the experiment with evidence or records why it was dropped. A git commit and MORNING.md make the changes inspectable in daylight.
The Guide, Organized as Six Setup Stages
Duncan's original guide has the exact prompts. These are the jobs they perform, in order; the early stages are manual on purpose.
| Stage | Output | Check before continuing |
|---|---|---|
| 1. Interview and define quality | AUDIENCE.md, RUBRIC.md, benchmarks/, initial PLAYBOOK.md, GUARDRAILS.md, FEEDBACK.md, and focus-axis files. | Can you tell a poor example from an excellent one using the written rubric? |
| 2. Build the run record | CLAUDE.md, runs/, ledger.jsonl, MORNING.md, and a retro prompt. | Are hard limits immutable and every change reversible? |
| 3. Create the critic | A separate critic prompt writing critique.md and score.json after benchmark comparison. | Does it cite visible evidence and score bad work harshly? |
| 4. Run one night by hand | Plan, artifact, critique, limited fixes, retro, ledger entry, commit, and morning report. | Did the entire loop finish inside time and cost caps without external action? |
| 5. Compress the report | A three-line MORNING.md: score, change, and what you need to review. | Can a person understand what happened without reading a long transcript? |
| 6. Schedule only after two clean manual runs | A local runner and scheduler; Duncan's example uses macOS launchd. | Confirm paths, permissions, power state, budget, and a way to stop the job. |
Claude Code's current non-interactive documentation explains claude -p, CLI options, and unattended permission behavior. Check those docs and settings precedence before copying the guide's runner: tool flags and defaults can change, and a scheduled process may load more project context and integrations than expected.
What the Three Web-Page Nights Actually Showed
| Run | Duncan's reported score | How to read it |
|---|---|---|
| Night 1, October 3 | 37 to 43 to 45 on an older 60-point rubric | He disliked the page despite the score. Do not compare this scale with later runs. |
| Hand-built benchmark | 72/90 after earlier iterations | A comparison page built by Duncan, not a public audience preference test. |
| Night 2, October 4 | 56 to 64/90 across eight rounds | It did not beat the 72-point bar; a losing night was recorded. |
| Night 3, October 5 | 73 to 77/90 across five scores | His critic judged the final page above the 72-point benchmark. The run took about an hour and used five images. |
Those numbers come from Duncan's guide. He later added a tenth category, making future scores out of 100. The observed improvement is within his own evaluation process; it does not yet show that visitors preferred the page, that customers converted, or that every later run improves. One of the three nights missed the bar.
Where He Applies the Loop, and What Is Still New
- Web pages: the clearest worked example. The agent researches a topic, writes a brief, builds the page and assets, captures screenshots, runs critique/fix rounds, then updates its rulebook.
- Instagram Reels: an early restyle lost to the original; a later staircase animation beat a simpler ladder-style visual in his internal comparison. The lesson was about making the idea clearer, not merely adding effects.
- Substack writing: he is beginning to train the system on his stories, opinions, and feedback. The written guide explicitly says this loop is new and has no results yet.
- Long-form video: discussed as a next application, not a proven overnight production line. Each format needs its own quality standards and review process.
What to Lock Down Before It Runs While You Sleep
- Keep authority narrow. Let the agent draft inside a dedicated folder. Do not grant publishing, sending, purchasing, deployment, or account-change authority to the overnight run. Treat morning review as a required handoff.
- Use real containment. Duncan notes that a broadly allowed shell can sidestep text-based edit deny rules. Prompt instructions and file deny-lists are not a security boundary. Use a disposable or isolated workspace, minimal credentials, restricted network/tools, and OS-level limits where appropriate.
- Cap time and spend outside the prompt. Put hard limits on runtime, model usage, and paid media generation. Stop on errors instead of retrying indefinitely.
- Preserve evidence. Keep benchmarks and hard limits human-owned. Log input, output, rubric version, score, changed rule, cost, and commit. Only compare runs under the same rubric or re-score the earlier one.
- Validate with people. Periodically blind-review samples, check the actual audience or business outcome, and reject rules that improve the critic's score while making the work worse.
The first pilot can be entirely manual. If it cannot produce a useful artifact, an honest critique, and a readable morning report twice in a row, scheduling it only automates the confusion.
Watch by Topic
| Time | Discussion |
|---|---|
| 00:00 | What Duncan means by a self-improving agent |
| 01:13 | Karpathy's autoresearch inspiration |
| 04:01 | The overnight web-page workflow |
| 05:28 | Training the critic's taste with examples |
| 10:42 | Rules the agent writes for its next run |
| 11:05 | Short-video experiments |
| 13:01 | Human feedback file |
| 14:03 | Substack writing loop |
A Business Idea Built From the Loop
The sellable version is not "an AI team that publishes while you sleep." It is a quality-improvement service for one repeated deliverable, with an owner-approved rubric and a human sign-off. Use this prompt to test whether a buyer would pay for that narrower outcome.
Test an overnight quality-loop service
A seven-day manual pilot before any scheduled agent.
Sources and Further Reading
- Creator: Duncan Rogoff's video and his full written setup guide. These are the primary sources for the workflow, six prompts, and reported scores.
- Inspiration: Karpathy's autoresearch repository, which uses objective training experiments. Duncan's creative-content judge is an adaptation, not the same experiment.
- Current tool behavior: Anthropic's non-interactive Claude Code guide and settings reference. Consult these before scripting an unattended run.