Startup Strategy

Max Hodak: Average Is Not Good Enough

Direct Answer

Max Hodak's central argument is that a startup's speed comes from its operating infrastructure, not only from the brilliance of its product team. Strategy becomes real through the systems that let employee 17 buy a piece of equipment, show what an experiment actually costs, move a candidate through a rigorous hiring funnel, surface performance problems early, and keep decisions attached to a human owner.

The talk is especially useful because Hodak rejects the usual promise of a universal founder checklist. Science Corporation builds medical devices, semiconductors, biological systems, and clinical programs. Its exact processes will not fit a two-person SaaS company. The transferable lesson is narrower: find the repeated delays that prevent your company from learning, then design an operating system that removes those delays without hiding cost, risk, or accountability.

The practical model: allocate resources deliberately, make routine action fast, instrument the result, review the evidence, and keep consequential judgment with a named person. Average processes produce average information. Better infrastructure increases the rate and quality of learning.

Watch Max Hodak at Startup School 2026

Video credit: Y Combinator and Max Hodak. Hodak is co-founder and CEO of Science Corporation. Product and clinical claims below are checked against Science's first-party material, ClinicalTrials.gov, and the peer-reviewed PRIMAvera paper. Internal costs, hiring conversion rates, and company processes remain Hodak's account unless stated otherwise.

What Science's Retinal Implant Has Demonstrated

Hodak opens with the product that gives the operational discussion its stakes. PRIMA is a visual prosthesis for people who have lost central vision because of geographic atrophy from age-related macular degeneration. A small photovoltaic array is implanted under the retina. Camera glasses capture the scene and project a processed pattern of near-infrared light onto the implant. The array converts that pattern into electrical stimulation for surviving retinal neurons, bypassing photoreceptors damaged by disease.

The pivotal PRIMAvera study was an open-label, multicenter, prospective, single-group study with 38 participants. Its peer-reviewed New England Journal of Medicine paper reported that the system restored useful central visual function for many participants, including meaningful visual-acuity gains and reading tasks. Science says 80% of participants demonstrated a clinically meaningful improvement. The device supplements remaining peripheral vision; it does not recreate normal color, field of view, or natural visual acuity.

Hodak says one participant used the system to finish a roughly 300-page novel and sent the book to the company. That is a powerful first-person outcome, but it is an anecdote rather than the trial's primary endpoint. Science's March 2026 update said it had submitted regulatory applications in Europe and the United States and was funding commercialization. Availability and approval status can change, so patients should use current regulator and clinical sources rather than this article for medical decisions.

The Claim Ledger: Evidence, Internal Data, or Founder Doctrine?

Claim from the talkEvidence statusResponsible interpretation
PRIMA restored useful central visual function in people with geographic atrophy.Peer-reviewed clinical evidence. The 38-person PRIMAvera study was open label and single group, with baseline comparison.A meaningful result for a specific patient population, not normal sight or a universal treatment for blindness.
A participant read a 300-page novel.Founder and patient anecdote. Consistent with reported reading capability, but not the study's primary outcome.Use it to understand lived value, not to estimate the average result.
A delayed $3,000 purchase can cost more than the saving.Operating example. The exact weekly burn and delay cost depend on the company.Compare purchase savings with the loaded cost of blocked people, equipment, revenue, and experiment time.
One Science wafer protocol cost about $40,000.Internal cost example. Not independently audited here.Shared materials and equipment must be allocated to experiments or teams will optimize against an imaginary zero price.
Science's application vote moves candidates to a phone screen quickly.Internal funnel data and process description.Measure stage conversion, time, quality, and adverse impact in your own funnel before copying the targets.
Eigen Reviews reveal who a company would hire again.Experimental internal method. No independent validation was presented.Treat graph-weighted peer ratings as one noisy signal with strong safeguards, never an automated employment decision.
Rate of iteration separates startups that win from those that fail.Founder doctrine with a plausible mechanism.Speed matters only when each loop produces trustworthy information and the team can act on it safely.

The Company Operating System

A product strategy says what the company wants to make true. A company operating system determines whether people can repeatedly do the work required to learn if that strategy is correct. Hodak's examples form seven connected layers.

LayerQuestion it must answerFailure signal
Resource allocationWhat are we willing to spend, and on which outcome?Every purchase becomes a fresh executive debate.
ProcurementHow does an authorized person obtain what the work needs?Approvals save hundreds while delays burn thousands.
Experiment costingWhat did this learning loop really consume?Shared materials make expensive experiments look free.
HiringHow do we identify job-relevant ability quickly and consistently?Founder instinct becomes a bottleneck or similarity filter.
Performance feedbackHow do people learn what is working before an annual review?Known problems remain unspoken until they are expensive.
Quality and safetyWhat evidence must exist before work moves forward?Speed creates undocumented or unsafe outcomes.
Decision ownershipWho interprets the evidence and accepts the consequence?Advice, committees, or software become substitutes for accountability.

The layers must connect. Fast procurement without experiment attribution creates uncontrolled spend. Strong hiring without feedback allows drift. Rapid iteration without safety gates can produce faster harm. An operating system is useful only when it increases speed and preserves visibility.

Procurement Creates Speed Before It Creates Savings

Hodak's test is wonderfully concrete: how does the seventeenth employee buy a $3,000 power supply? A founder can use a credit card. A growing company introduces permissions, budgets, vendor records, insurance requirements, purchase orders, invoices, and accounting. Each control has a purpose, but unmanaged controls can create a queue longer than the experiment itself.

The key design move is to put the economic decision earlier. Leadership sets budgets, permitted categories, risk tiers, and approval ceilings. Inside those boundaries, routine purchasing should be fast. Review unusual vendors, large commitments, security exposure, regulated materials, and irreversible contracts. Do not force a fresh strategy meeting for every ordinary component.

Cost of delay

blocked people x loaded hourly cost x blocked hours
+ idle equipment or facility cost
+ expiring materials and rescheduling
+ missed revenue or milestone value
+ risk created by the delay
= total delay cost

Decision rule
If delay cost is greater than purchase saving,
the cheaper purchase is economically more expensive.

Measure procurement like an operational system: median request-to-order time, 90th-percentile time, percentage of requests returned for missing information, emergency purchases, budget variance, and delay cost. The goal is not maximum spending freedom. It is the shortest accountable path from a legitimate need to the work.

An Experiment Is Not Free Because the Invoice Is Shared

Deep-tech teams often buy gases, resins, media, wafers, animal work, external assays, and equipment in bulk. When those costs sit in a shared account, each experiment feels free to the person choosing it. Hodak says Science connected purchasing data to its manufacturing and laboratory records and discovered that one wafer iteration could cost roughly $40,000.

Exact allocation will never be perfect. It does not need to be. A consistent approximation is better than a precise-looking zero. Use the same rules across candidate protocols, document assumptions, and show a range when uncertainty is material.

Cost componentAllocation methodCommon omission
Direct materialsActual units consumed plus expected scrapFailed runs and expired stock
LaborLoaded hourly cost by role and timePreparation, cleanup, analysis, and supervision
EquipmentInternal hourly rate, lease cost, or depreciation plus maintenanceSetup and idle time reserved for the run
Facilities and shared consumablesPer run, machine hour, bench hour, or documented percentageGases, utilities, cleanroom, storage, and waste
Quality and regulatoryReview hours, documentation, validation, and external feesEvidence generation after the technical work
Failure and reworkBase cost multiplied by expected failure probabilityThe first run is silently assumed to work
Lead timeCost of delay for the critical pathA low-price vendor that moves the milestone
Better metric: optimize cost per accepted learning outcome, not cost per run. A $10,000 experiment that answers the decision is cheaper than five $3,000 experiments that leave the team uncertain.

Build a Hiring System Before Founder Instinct Becomes the Queue

Early startups recruit from the network and technical scene that produced the company. That source can be excellent, but it eventually runs out and can reproduce the same backgrounds. Hodak says Science built a four-stage funnel: distributed application voting, a phone screen, a job-relevant homework exercise, and a full interview. The system is designed to make broad participation possible without sending every decision through the founder.

He describes three traits: judgment, horsepower, and agency. The second term is memorable but imprecise. A safer scorecard turns all three into observable behavior.

CriterionObservable definitionEvidence prompt
JudgmentMakes evidence-based tradeoffs under uncertainty and names what would change the decision."Tell us about a decision with incomplete data. What alternatives did you reject, and what happened?"
Learning capacityAcquires a difficult concept, transfers it to a new problem, and corrects errors.Give a short unfamiliar brief, then ask the candidate to build and explain a solution.
AgencyOwns an outcome, finds constraints, communicates risk, and follows through."Show a result you moved without formal authority. What did you personally do?"
Role craftPerforms the actual work to the required standard.Use a paid, time-bounded work sample with a clear rubric and realistic tools.

Hodak prefers work samples with a high ceiling, numerical scoring, and a moving frontier. AI use can be allowed when the job itself permits AI. The purpose is not to create an "AI-proof" puzzle. It is to observe problem framing, tool choice, verification, communication, and the accepted result. A candidate who uses an agent well may be demonstrating the job more faithfully than one forced to hide it.

Structured interviews and work samples still need governance. The U.S. Equal Employment Opportunity Commission says selection procedures should be job related, validated for their purpose, and checked for discriminatory impact. The U.S. Office of Personnel Management's structured-interview guidance recommends consistent questions, anchored rating scales, trained interviewers, and job analysis. Other jurisdictions impose different duties, so employers need local advice.

Eigen Reviews Are an Experiment, Not a Default

Hodak criticizes annual 360-degree reviews because they are disruptive, delayed, and often restate problems everyone already knew. Science instead asks employees every four to six weeks whether, knowing what they know now, they would vote to hire a colleague again. Its internal system weights responses through a graph inspired by eigenvector centrality: ratings from people who are themselves trusted carry more weight.

The idea is clever and the risk is substantial. A network score can confuse visibility with value, amplify dominant groups, penalize dissent, hide retaliation, and make social popularity look mathematical. Detecting cliques with repeated graph runs or dropout does not remove the underlying employment and power issues.

Required safeguardWhy it matters
Behavioral criteriaReview job outcomes and conduct, not whether someone feels culturally familiar.
Multiple evidence typesCombine work output, manager context, peer feedback, self-reflection, and role expectations.
Human calibrationInvestigate disagreement and context instead of accepting a composite score.
Right to respondEmployees must be able to inspect material concerns, correct errors, and appeal.
Bias and adverse-impact testingCheck whether the process systematically disadvantages protected or less powerful groups.
Privacy and retention controlsLimit who can see raw feedback, how long it remains, and how it can be reused.
No automated adverse actionA graph score must not decide termination, promotion, compensation, or access by itself.

The more transferable alternative is continuous, specific feedback attached to observed work. Ask what outcome was expected, what happened, what evidence supports the conclusion, what support is missing, and what changes next. Shorter loops are valuable; opaque social ranking is optional.

Iteration Speed Is an Infrastructure Property

Hodak argues that a company learning weekly will separate from a competitor learning monthly. The compounding intuition is useful, but speed is not the number of tickets closed or experiments started. An iteration counts only when it changes what the team knows and improves the next decision.

Accepted learning loop

observation
-> explicit assumption
-> smallest reversible action
-> success and failure thresholds
-> controlled execution
-> independent result check
-> decision and owner
-> updated operating rule

Procurement shortens the wait before execution. Experiment costing helps choose the right test. Hiring provides the required skill. Quality systems define acceptable evidence. Feedback fixes recurring failure. The infrastructure is not overhead around iteration; it is what makes responsible iteration possible.

Track median idea-to-result time, percentage of experiments with predeclared thresholds, cost per accepted learning outcome, repeat-failure rate, time from evidence to decision, and rollback or corrective-action time. A team can then distinguish genuine learning speed from frantic activity.

You Cannot Delegate Founder Judgment

Hodak's most important warning is that founders cannot outsource judgment. Advisers can supply experience. Employees can own domains. Investors can identify patterns. AI can retrieve evidence, model scenarios, and challenge assumptions. None of them accepts the founder's accountability for the final company-level tradeoff.

This is not an argument for ignoring expertise. It is an argument against averaging advice until every decision resembles the market. Successful startups are long-tail outcomes. Copying an average process can be useful for commodity work; copying an average strategic opinion can erase the reason the company exists.

Hodak says action produces information. Once a team takes a reversible step, the world reveals constraints that discussion could not. The mature version of that rule is: act where the downside is bounded, instrument the action, and keep irreversible decisions behind stronger evidence and human review.

AI boundary: let an agent assemble the decision packet, identify missing evidence, model alternatives, and record the outcome. Do not let it silently become the accountable executive for hiring, medical safety, regulatory claims, major spending, or strategy.

When to Build Internal Software and When to Buy It

Science built internal tools for purchasing, manufacturing data, recruiting, and reviews because its workflows did not fit ordinary software. Hodak argues that AI coding agents now lower the cost of creating tailored infrastructure. That expands the build option, but it does not remove maintenance, security, data governance, or continuity costs.

Build whenBuy when
The workflow encodes a real operating advantage.The function is commodity, regulated, and already solved well.
Existing software creates measured delay or data fragmentation.The problem is occasional annoyance rather than repeated cost.
The process is understood well enough to encode and test.The team is still discovering what the process should be.
A named team can own security, uptime, migrations, and support.No one can maintain the system after the first builder leaves.
Integration and data ownership justify the lifecycle cost.Payroll, tax, identity, payments, or compliance exposure dominate.

Start with a manual workflow and measure it. Build the smallest internal interface that removes the proven bottleneck. Preserve source data, logs, permissions, tests, export, and an exit path. Agent-generated code can make version one cheaper; it cannot make ownership free.

A 30-Day Startup Operating-System Audit

Week 1: purchasing and delay

Trace ten recent requests from need to order. Record handoffs, waits, rework, approval value, and cost of delay. Set budget bands and risk tiers, then remove one redundant approval while keeping audit logs and escalation rules.

Week 2: experiment economics

Choose three representative experiments or product releases. Allocate direct materials, labor, equipment, shared costs, quality work, failure, and lead time. Compare cost per accepted learning outcome and decide which protocol should change.

Week 3: hiring and feedback

Define four job outcomes, create anchored criteria, and build one realistic paid work sample. Measure funnel conversion and time by stage. Replace one vague performance question with a specific expectation, observed evidence, next action, and review date.

Week 4: decision loops and internal tools

Map one recurring decision from signal to owner. Add success and failure thresholds, a human gate, a worklog, and a review cadence. Only then decide whether a lightweight internal tool would remove enough repeated friction to justify ownership.

Audit output: one latency map, one real experiment-cost model, one structured hiring scorecard, one shorter feedback loop, and one named decision owner. That is enough to improve the company OS without launching a bureaucracy project.

Copy-Ready Operating Worksheets

Experiment brief

Decision this experiment must inform:

Assumption:

Success threshold:
Failure threshold:

Direct materials:
Labor:
Equipment:
Shared facilities and consumables:
Quality / regulatory / external testing:
Expected failure and rework:
Lead-time cost:

Total expected cost:
Expected time to accepted result:

Owner:
Decision date:

Structured hiring scorecard

Role outcome:

Criterion 1: judgment
Observable behavior:
Evidence:
Anchored score (1-5):

Criterion 2: learning capacity
Observable behavior:
Evidence:
Anchored score (1-5):

Criterion 3: agency
Observable behavior:
Evidence:
Anchored score (1-5):

Criterion 4: role craft
Work-sample result:
Verification quality:
Anchored score (1-5):

Accommodation offered:
Conflicts or bias risks:
Decision and rationale:
Independent reviewer:

Founder decision record

Decision:
Accountable owner:

Evidence for:
Evidence against:
Unknowns:

Reversible or irreversible:
Maximum acceptable downside:

Options considered:
Chosen action:

What would change our mind:
Review date:
Observed result:
Operating rule updated:

Video Chapters

TimeTopic
00:00Infrastructure at startups
01:02Science's retinal implant
02:41The hidden infrastructure of a startup
03:25How employee 17 buys things
06:15How infrastructure creates speed
07:03What an experiment actually costs
09:27How the best startups hire
10:18Building a rigorous hiring process
13:26Judgment, learning capacity, and agency
16:17Rethinking performance reviews
18:53Rate of iteration separates outcomes
20:43You cannot delegate your judgment
22:47Action produces information
24:31The operating system of a company
25:10Audience Q&A

Bottom Line

Max Hodak's talk is not really about procurement software or a clever review algorithm. It is about the distance between a founder's intention and what the organization can do on Tuesday morning. That distance is filled by budgets, permissions, cost models, scorecards, quality evidence, feedback, and decision ownership.

"Average is not good enough" should not become permission for arbitrary processes or heroic founder instinct. The sharper lesson is to stop borrowing operating assumptions without testing them. Measure where learning stalls. Build the smallest system that removes the constraint. Keep costs visible, employment decisions fair, medical claims precise, and accountability human.

A startup does not need enterprise bureaucracy. It needs an operating system proportionate to its work. When that system is good, the company buys faster, learns what experiments cost, identifies talent with better evidence, surfaces problems earlier, and gives founders higher-quality information for the judgments only they can own.

Sources and Link Map

Common questions

What does Max Hodak mean by startup infrastructure?
He means the operating systems that turn strategy into repeated action: purchasing, budgeting, experiment costing, recruiting, performance feedback, quality, safety, and internal software. These systems determine how quickly a team can learn without losing control of money, risk, or accountability.
What did Science Corporation's PRIMA retinal implant demonstrate?
In a 38-participant, open-label clinical study involving people with geographic atrophy caused by age-related macular degeneration, the PRIMA system restored useful central visual function for many participants. The peer-reviewed study reported meaningful gains in visual acuity and the ability to perform tasks such as reading letters, numbers, and words. It does not restore normal vision and is designed for a specific form of central vision loss.
Is PRIMA already generally available?
The latest first-party information verified for this article says Science submitted applications in Europe and the United States and was preparing for commercialization. Availability, eligibility, risks, and regulatory status must be checked with Science, trial investigators, and the relevant regulator. This article is not medical advice.
How should a startup calculate the cost of delay?
Start with the loaded cost of the people and equipment blocked by the delay, then add missed revenue, expiring materials, rescheduling, and opportunity cost. Compare that figure with the proposed purchase saving. A cheaper component is not cheaper if a week of delay costs more than the discount.
How should a deep-tech startup cost an experiment?
Include direct materials, labor, equipment time or depreciation, shared consumables, facilities, quality and regulatory work, external testing, expected failure and rework, and the cost of lead time. Allocate shared costs with a documented rule so teams can compare protocols consistently.
What are judgment, horsepower, and agency in Hodak's hiring model?
This article translates them into less ambiguous criteria: judgment is making evidence-based tradeoffs under uncertainty; learning capacity is acquiring and transferring difficult knowledge; and agency is owning an outcome through constraints and follow-through. Each criterion should be tied to observable job behavior, not intuition or personality similarity.
Are algorithmic peer reviews such as Eigen Reviews a good idea?
They are an interesting internal experiment, not an established universal best practice. Network-weighted ratings can amplify popularity, power, retaliation, and group bias. They should never be the sole basis for employment action and require transparent criteria, calibrated human review, an appeal route, adverse-impact testing, privacy controls, and local employment-law advice.
Should startups build their own internal software?
Build when a repeated workflow is strategically differentiating, current tools impose measurable delay, requirements are stable enough to encode, and the company can maintain the system. Buy commodity functions such as payroll, tax, identity, and standard accounting unless there is an exceptional reason to own them.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call