AI Web Design

How to Design in the Agent Era: A Human-in-the-Loop Playbook

Direct Answer

Design in the agent era is not prompt-and-ship. It is a loop in which agents make options cheap and humans make decisions expensive. Use agents to expand the possibility space, apply repetitive changes, connect real content, and turn approved designs into code. Use a visual canvas to compare alternatives. Then let a person remove generic patterns, protect trust, resolve stakeholder constraints, and verify the live result.

That is the durable lesson in Y Combinator's conversation with Paper founder Stephen Haney. Faster software production does not lower the need for design. It raises the cost of having no opinion. When every team can generate a polished-looking page, differentiation comes from what the interface chooses to emphasize, remove, prove, and make unmistakably its own.

The operating loop: brief the problem, expand into alternatives, compare them spatially, subtract everything without a job, then verify the real product.

Watch How to Design in the Agent Era

Video and editorial credit: Y Combinator. The episode features Paper founder Stephen Haney with YC General Partner Aaron Epstein. The live reviews, product demonstrations, opinions, and company anecdotes belong to the speakers. JQ AI SYSTEMS adds the implementation framework, verification checks, and evidence labels below.

Paper's current official site confirms the product architecture demonstrated in the episode: a canvas built on web standards, a desktop app with MCP connectivity, design-to-code and code-to-design workflows, and agent access to tokens, styles, components, and real data. Features are still evolving; use the current roadmap rather than assuming every previewed capability is generally available.

What Agents Actually Change About Design

Agents compress production. They can generate twenty directions, translate a page, resize a campaign, apply tokens, create responsive variants, pull data, and write implementation code. That moves the bottleneck upstream. The difficult question is no longer only "Can we build this?" It is "Which version deserves to exist, for whom, and what evidence makes it trustworthy?"

Design responsibilityAgent contributionHuman responsibility
Problem framingSummarize research, organize constraints, and surface contradictions.Choose the audience, business objective, risk tolerance, and decision criteria.
ExplorationGenerate bounded variations across layout, type, color, content, or interaction.Judge which directions are relevant rather than merely attractive.
ProductionResize, translate, populate real content, create states, and implement approved patterns.Protect coherence, accessibility, brand meaning, and edge-case priorities.
ReviewRun structured audits and show before-and-after alternatives.Resolve tradeoffs, challenge assumptions, and approve the final direction.
ShippingGenerate code, tests, screenshots, and handoff artifacts.Verify the real product and own the consequence of release.

Haney's strongest point is organizational: the pixels are an output of design, not its full definition. Product design includes competitive analysis, requirements, stakeholder alignment, value-proposition decisions, and choosing what not to build. Models can help with each, but speed does not grant them authority.

Why Paper Uses HTML, CSS, and MCP

Paper's technical bet is that the visual canvas and the shipped product should speak the same language. Its official product page says the canvas is built on HTML and CSS. The desktop application exposes an MCP connection so supported agents can read and write the canvas, sync design tokens and components, fetch content, and move between design and code.

In the demonstration, the agent sends HTML rather than translating through a proprietary drawing format. Designers can copy work as React or Tailwind, and Paper's current download page lists connections for Claude Code, Codex, Cursor, Copilot, and OpenCode. Paper Snapshot can bring a live web section back into the canvas as editable layers. This creates a bidirectional loop instead of a one-way handoff.

LayerPurposeControl to preserve
CanvasKeep alternatives, references, annotations, and direct manipulation visible.The designer can select, move, rewrite, and delete without re-prompting.
Web standardsRepresent layout in a language agents and browsers already understand.Inspect the generated DOM, CSS, responsive behavior, and semantics.
MCP connectionLet agents exchange structured context and actions with Paper.Limit tools, repositories, data sources, and write permissions to the task.
CodebaseKeep production components, tokens, and behavior close to the design source.Use version control, review diffs, run tests, and prevent secrets entering design context.
Live productReveal actual content, device behavior, performance, and user constraints.The shipped interface, not the canvas, is the final evidence.

Paper says this architecture improves speed, accuracy, and token use. Treat those as vendor claims until reproduced in your own files. The more important architectural benefit is inspectability: a design element is closer to the material of the web, and a person can still manipulate it visually.

Paper also publishes Paper Shaders, an open-source, zero-dependency React shader collection with installable effects and image filters. The library makes sophisticated motion and texture easier to reach. Accessibility, performance, brand fit, and restraint remain the implementer's responsibility.

The Five-Stage Human-Agent Design Loop

1. Brief the decision, not the picture

Start with the audience, task, proof, emotional register, trust requirement, content priority, and constraints. "Make a modern landing page" invites the average of the training data. "Help a finance lead verify cross-border payment coverage before creating an account" gives the interface a job.

2. Expand one variable at a time

Ask for alternatives, but control the axis. Generate five information hierarchies with the same visual system, then five type directions using the chosen hierarchy. If content, layout, typography, color, imagery, and motion all change at once, comparison becomes theater.

3. Compare spatially

A chat log hides discarded possibilities. A canvas keeps several futures visible. Place alternatives side by side with the same real content and viewport. Annotate what each version makes clear, what it hides, and the user decision it supports. Pick with declared criteria, not with the last output's novelty.

4. Subtract before polishing

Haney repeatedly improves the reviewed pages by reducing weight, sizes, cards, badges, text, and decorative widgets. Deletion is a design operation. If an element does not explain the product, create trust, establish hierarchy, support an action, or express a distinctive brand idea, remove it and see whether anything is lost.

5. Verify in the real medium

Push the approved direction into code, use production content, and inspect real breakpoints and states. Verify keyboard navigation, focus, semantics, contrast, text resizing, slow loading, empty data, errors, localization, analytics, and performance. A canvas can approve intent; only the running product can prove behavior.

The AI Design Smell Checklist

The episode identifies recurring visual patterns in generated sites. These are smells, not bans. A pill, gradient, or card can be exactly right. The problem is using the pattern because the model offered it, not because the interface needs it.

SmellWhy it weakens the pageReview action
Everything is boldHierarchy collapses when every line asks for maximum attention.Use the lightest weight that preserves hierarchy; reserve bold for real emphasis.
Five to eight type sizesThe page feels assembled from widgets rather than designed as a system.Try three primary sizes per section or view before adding exceptions.
Cards everywhereEvery idea becomes an isolated box, adding borders and gaps without clarifying relationships.Remove containers; rebuild hierarchy with alignment, spacing, rules, and typography.
Pills, badges, and tiny iconsDecorative labels imitate product UI and consume attention without evidence.Keep only labels that change comprehension, status, or action.
All-caps micro-kickersWide letter spacing and tiny uppercase labels have become a generic AI-era signature.Use plain language or remove the pre-heading if the main heading already works.
Purple glows and gradientsThe style can signal "AI startup" before it communicates a specific brand.Choose color from audience, category, history, and differentiation; not model default.
Decorative metrics and widgetsNumbers, mini charts, and corner ornaments can look authoritative while saying nothing.Require a real source, meaning, and user decision for every metric.
Automatic light and dark modeA generated feature can consume design and QA time without helping the launch goal.Ship the mode the audience needs; add the second only with a reason and full state review.
Too much copyAgents explain every thought, pushing proof and action below the fold.Keep one promise, one proof path, and one primary action in the first viewport.
Effects without a roleMotion, shaders, and texture become decoration rather than communication.Connect the effect to state, hierarchy, brand behavior, or a meaningful interaction.

The quickest audit is to ask of every component: what changes for the user if this disappears? If the honest answer is "the page looks less generated," deletion is probably an improvement.

What the Three Live Reviews Teach

Legion Health: clarity is not the same as finish

The reviewers can understand the psychiatric-care offer quickly, which is a real success. Their criticism targets the generic promotional pill, small spacing gaps, and weak brand finish. The lesson is to preserve the clear value proposition while making the trust layer deliberate. In health, reassurance, eligibility, privacy, clinician credibility, and the next step matter more than ornamental novelty.

Sytex: the visual world must match the operating world

Sytex is about field infrastructure operations, yet the page initially reads like a generalized software dashboard with bold type, purple glow, and abstract UI. The most useful feedback is not "remove purple." It is "make the user outcome and field context obvious." Real equipment, crews, process states, and before-and-after operational proof can connect the software to the environment where it earns value.

Moreta: energy and trust must coexist

Moreta's global payment page has energetic flags and an international feel, but payment products carry a high trust burden. Haney recommends bringing the useful "how it works" content higher and building a more professional foundation without erasing the brand's energy. The right move is not sterile minimalism. It is structured excitement: clear coverage, security, fees, steps, and proof, with expressive elements supporting rather than replacing them.

Trust-sensitive rule: as the consequence of error rises, visual confidence must be backed by more explicit proof, clearer process, and less decorative ambiguity.

Can AI Learn Taste?

Haney expects models to improve at tactics: typography weight, fewer arbitrary elements, better spacing, and more thoughtful first drafts. He is more skeptical that they can replace the organizational role of design. Epstein adds the harder question: can experts verbalize the implicit criteria behind their reactions well enough to teach a model?

The useful answer is to decompose taste instead of mystifying it. Some parts are teachable rules. Some are pattern recognition from a body of work. Some depend on product strategy, market position, cultural timing, risk, stakeholder history, and a particular team's values. Agents can score declared criteria; they cannot choose the criteria without inheriting someone else's priorities.

Taste also moves. Once a successful style saturates training data and generated output, using it well is no longer enough to differentiate. The episode's Linear-style purple gradient example captures this dynamic: a good design language becomes generic through imitation. Humans remain responsible for noticing when the baseline has moved.

How Paper Uses Agents Without Lowering the Bar

Paper's internal practice is a useful counterweight to the product hype. Haney says the roughly twelve-person team uses coding agents extensively but still has humans read every line of core product code, often more than once. A high-performance design tool needs precision, speed, and reliability that the team does not delegate blindly.

The company uses more agent leverage around the core: its brand designer built and shipped the marketing site from Paper designs with agent help, and the team uses agents for videos, campaign assets, coding, and pull-request review. That is a risk-based boundary. Repetitive and reversible surfaces get more autonomy; the editor's performance-critical core gets deeper human review.

The broader operating lesson is not "move slowly." It is to keep a small, capable team and use agents where they multiply the team without erasing understanding. Shipping more features is not the only measure of product progress. Reliability, care, and whether users want the product to improve are strategic assets.

Claim Ledger

Claim or impressionGrounded readingStatus
Paper is one of the fastest-growing design tools since Figma.A strong host description. Paper's official site documents growth and a $34M Series A, but the comparative ranking is not independently established here.Promotional claim
HTML/CSS makes agent flows faster, cheaper, and more accurate.The shared language can reduce translation work. Paper makes benchmark claims, but teams should reproduce them on their own designs and agents.Vendor claim with plausible mechanism
Paper Shaders is open source.The official shader site provides installable React packages and source-linked usage for the collection.Documented
Great companies have exceptional design.A persuasive design philosophy, not a universal causal finding. Distribution, timing, product value, operations, and market structure also matter.Opinion
AI will not replace human design decisions.A forecast. Current evidence supports keeping humans in problem framing, stakeholder judgment, trust, and approval, but future capability is uncertain.Prediction
Paper's brand designer shipped the company site in one week.An internal anecdote described by the founder; useful as a workflow example, not a controlled productivity benchmark.Creator report
Paper reviews every line of core code with humans.The founder's description of current team practice in the episode.Creator report
Agent-generated pages share recognizable visual tells.The live reviews supply examples, but the patterns are heuristics. They should not be used to accuse a team of AI use or to ban styles categorically.Review heuristic

Three Copy-Ready Design Prompts

1. Expand without changing everything

Create six alternatives for this page using the exact same content,
brand tokens, components, and viewport.

Change only the information hierarchy and layout. Keep typography,
color, imagery, and interaction behavior fixed.

For each alternative, label:
- primary user task
- first thing noticed
- proof surfaced above the fold
- tradeoff introduced

Place all six side by side. Do not choose a winner.

2. Run an anti-slop subtraction pass

Audit this design for elements that exist without a communication job.
Check: font weights, number of type sizes, cards, pills, badges, icons,
all-caps kickers, gradients, glows, decorative metrics, duplicate copy,
automatic dark mode, and effects without meaning.

For every flagged element, state its current job. If no job can be
identified, create a version with it removed. Preserve the value
proposition, required proof, primary action, and distinctive brand idea.

Return before and after side by side. Do not make production changes.

3. Turn an approved design into a verification contract

Implement the approved frame using the existing production components
and tokens. Do not invent copy, data, icons, or interactions.

Verify at 390, 768, 1024, and 1440 px. Check text overflow, keyboard
navigation, focus visibility, semantics, contrast, loading, empty,
error, success, and reduced-motion states.

Return:
1. changed files and diff,
2. screenshots at every viewport,
3. automated checks and results,
4. visual differences from the approved frame,
5. unresolved decisions.

Do not deploy. Stop if the design conflicts with an existing component,
accessibility requirement, or product behavior.

A 45-Minute Agent-Era Design Workshop

  1. Minutes 0-5: define the decision. Name the user, task, trust burden, evidence, primary action, and one thing the page must feel unlike.
  2. Minutes 5-12: collect reality. Add production copy, data, screenshots, current components, constraints, and three relevant references.
  3. Minutes 12-20: expand one axis. Generate six information hierarchies with everything else fixed.
  4. Minutes 20-27: compare visibly. Score clarity, proof, differentiation, trust, and implementation risk. Keep the discarded versions visible.
  5. Minutes 27-33: subtract. Remove unsupported cards, labels, icons, weights, sizes, effects, and copy.
  6. Minutes 33-38: restore the human idea. Add one intentional brand behavior, image system, interaction, or narrative device that the default generation missed.
  7. Minutes 38-45: write the contract. Define responsive states, accessibility checks, content ownership, implementation evidence, and the person who can approve shipping.
Pass condition: another team member should be able to explain why the chosen direction won, which alternatives were rejected, what the agent did, and what still requires human approval.

Video Chapters

TimeChapterDesign lens
00:00Great companies have great designWhy cheap execution raises the value of differentiation.
01:04What is Paper and why build it?Founder-market fit and a new design-tool architecture.
03:17What makes Paper agent-nativeHTML/CSS as the shared language for canvas and agent.
05:11Shaders, image generation, and brand designPowerful effects still need brand purpose and restraint.
11:32Design-to-code and the new agent stackMCP, React, Tailwind, Snapshot, and bidirectional work.
16:48Design review: Legion HealthKeep value clarity; improve finish and trust.
24:11How to avoid AI design slopReduce type weight, type sizes, and default patterns.
28:26Design review: SytexConnect abstract software to field operations and outcomes.
32:42Biggest tells of AI-generated designCards, pills, icons, gradients, kickers, and empty widgets.
37:39Design review: MoretaBalance expressive global energy with financial trust.
41:11Can AI learn taste?Tactical rules versus contextual judgment.
43:56How Paper uses agents internallyRisk-based autonomy and deep review for the core product.
46:52Building a community around PaperValues, product quality, and market alignment.
49:15Lessons from Stephen's first startupBottlenecks, user conversations, and simpler strategy.
53:30What is next for design and PaperFaster tooling, persistent human design work, and prototype feedback.

Bottom Line

Agents are excellent production multipliers. They make variation, translation, resizing, content population, responsive implementation, and code generation dramatically easier. Paper's interesting contribution is giving those agents a visual, HTML/CSS-backed canvas where a person can still compare, manipulate, annotate, and delete.

The resulting workflow is neither traditional mockup handoff nor autonomous prompt-and-ship. It is a connected loop. Humans define the problem and hold several possibilities in view. Agents perform the repetitive transformations. Humans narrow the field, protect meaning and trust, and approve the result. The codebase and live product provide the final evidence.

The biggest risk is not that every generated interface looks terrible. It is that every interface looks competent in the same way. A useful agent-era designer therefore needs two habits: explore more than one future, and delete more than the model expects. Speed creates the room. Care decides what fills it.

Sources and Useful Links

Common questions

What does agent-native design mean?
Agent-native design means the design environment is built so AI agents can inspect, create, and modify real design structures instead of only returning screenshots or chat artifacts. Paper uses an HTML/CSS canvas and an MCP connection so supported coding agents can work with the same visual surface as the designer.
Does agent-native design replace designers?
No. Agents can generate variations, resize, translate, connect content, and perform repetitive production work. Designers still frame the problem, compare directions, make tradeoffs, manage stakeholders, protect trust, and decide what the product should communicate.
What are the most common signs of AI-generated web design?
The episode highlights excessive bold type, too many font sizes, overused cards, pills and badges, tiny all-caps kickers, decorative icons, purple gradients and glows, empty widgets, automatic dark mode, and too much copy. None is inherently wrong; the smell comes from using them without a communication reason.
Is Paper a replacement for Figma?
Not as a universal claim. Paper is differentiated by an HTML/CSS canvas, agent connections, code export, live-site capture, and visual iteration. Figma remains deeply established for collaborative product design and mature design systems. Choose around the workflow and team, not the slogan.
Do designers need to learn to code in the agent era?
They do not need to become full-time engineers, but understanding layout, components, responsive behavior, tokens, Git, and verification makes collaboration with coding agents much stronger. Paper is designed to let direct visual manipulation and code-based agents share one workflow.
Can AI learn design taste?
Models can improve at tactical conventions and can imitate examples, but taste also includes product context, values, audience, risk, timing, organizational constraints, and stakeholder judgment. The episode treats those human decisions as the durable part of design.
What should an agent do in a design workflow?
Good agent tasks include generating bounded alternatives, applying real content, checking responsive states, translating and resizing, importing tokens, producing code, and running structured audits. The person should choose the direction, remove noise, resolve conflicting requirements, and approve production output.
How should AI-generated designs be verified before shipping?
Check value-proposition clarity, real content, responsive layouts, text overflow, keyboard and screen-reader behavior, contrast, loading and error states, trust cues, code quality, performance, analytics, and whether every visual element has a job. Review the live product, not only the canvas.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call