AI Policy

Daniel Kokotajlo's 70% AI Warning: What He Actually Said

Direct Answer

Daniel Kokotajlo is making an unusually severe forecast: he gives roughly a 70% probability that the development of advanced AI goes horribly wrong. He does not reduce that estimate to a neat 70% chance of human extinction. In the interview, he explicitly says extinction is one possible outcome inside a wider category that includes loss of control, AI takeover, and other major catastrophes.

He also gives 2029 as his current median for superintelligence, argues that most economically useful work could eventually be automated, and supports a US-China agreement to slow the frontier before automated AI research accelerates beyond human control. Those are forecasts and policy judgments, not established facts.

The useful way to read this conversation: separate the documented history from Kokotajlo's probability estimates, then separate both from the conditional futures in AI 2027 and the policy recommendation in AI 2040: Plan A.

Watch the Full Interview

Video credit: The Diary of a CEO, hosted by Steven Bartlett, with guest Daniel Kokotajlo. Follow Daniel through the show links on X and Substack. This article is an independent, source-checked companion to the interview.

Four Corrections Before the Big Claims

Popular shorthandMore accurate readingStatus
"70% chance AI causes extinction."Kokotajlo says roughly 70% that things go horribly wrong; extinction is one possible failure mode.Personal probability estimate
"He lost $2 million to speak."He refused the clause while expecting to lose roughly $2M in vested equity, but OpenAI changed course and he ultimately retained it.Documented event with later correction
"Superintelligence arrives in 2029."2029 is his current 50% point, with acknowledged probability on much later timelines.Forecast, not deadline
"AI 2040 predicts Plan A will happen."The authors explicitly call Plan A primarily a recommendation, not their best guess of what governments will do.Policy scenario

These distinctions do not make the concerns trivial. They make the argument testable. A dramatic claim becomes more useful when its probability, scope, time horizon, source, and conditions are visible.

Who Daniel Kokotajlo Is

Kokotajlo joined OpenAI in 2022 as a governance researcher focused on forecasting and scenario planning. He contributed to work around GPT-4 and left in 2024 after losing confidence that the company would act responsibly as capabilities increased. He now leads the nonprofit AI Futures Project.

The project's first major public scenario, AI 2027, was written by Kokotajlo, Eli Lifland, Thomas Larsen, Romeo Dean, and Scott Alexander. The site says it was informed by trend extrapolations, more than two dozen tabletop exercises, feedback from over 100 people, and separate forecasts on compute, timelines, takeoff, goals, and security.

That is meaningful work, but proximity is not omniscience. An insider can know organizational incentives and technical discussions better than the public while still being wrong about timelines, bottlenecks, economics, or geopolitics. Kokotajlo's background is a reason to examine the argument closely, not a substitute for examining it.

Why He Left OpenAI, and What Happened to the $2 Million

Kokotajlo says he resigned because he no longer trusted OpenAI to behave responsibly around increasingly capable systems. His June 2024 public thread says the company asked him to sign exit paperwork containing a non-disparagement clause and that the paperwork and communications indicated he would lose vested equity if he refused.

He refused because he wanted the freedom to criticize the company. The equity was commonly valued at roughly $2 million and represented most of his household net worth. Public reporting and Kokotajlo's own account support the central point: he accepted the possibility of a very large personal financial loss rather than sign.

Important ending to the story: OpenAI later changed its non-disparagement policy and said former employees would not lose vested equity for refusing the clause. Kokotajlo has since confirmed that he kept the equity. "He was prepared to lose $2 million" is accurate; "he permanently gave away $2 million" is not.

The episode remains relevant because frontier-lab employees may hold information the public cannot independently inspect. Financial penalties, broad confidentiality rules, and weak whistleblower protections can distort which concerns reach regulators and the public, regardless of whether a particular warning later proves correct.

Why His Median Is 2029

Kokotajlo defines superintelligence as AI better than the best humans across cognitive domains, while also operating faster, more cheaply, and eventually through robots in the physical world. His median estimate in the interview is 2029. A median is not a certainty: it means he assigns equal probability to arrival before and after that point.

His argument is a chain, and every link matters:

  1. AI systems continue improving at coding and research. The leading labs increasingly use AI to accelerate their own development work.
  2. AI research becomes substantially automated. Models run experiments, write and debug code, analyze results, and propose the next experiments with declining human input.
  3. Research acceleration feeds back into capability. Better AI helps produce better AI, shortening the cycle again.
  4. Competition reduces willingness to pause. Companies and states fear that slowing down gives a rival decisive strategic power.
  5. Alignment, security, and governance lag capability. The systems become harder to evaluate while the incentives to deploy them increase.

The central uncertainty is step two. Current systems can accelerate portions of software and research, but automating the full scientific loop requires reliability over long horizons, experimental judgment, secure infrastructure, tacit knowledge, and progress through physical bottlenecks. Small changes in those assumptions can shift the date by years or decades.

AI 2027 itself now carries a useful clarification: 2027 was the authors' modal year at publication, not a shared median deadline, and their medians ranged later. The project maintains a changelog and updates its forecasts. That transparency is a strength, and it also shows why the scenario title should not be mistaken for a countdown clock.

What the 70% Estimate Really Means

Around 42 minutes into the interview, Kokotajlo corrects the strongest version of the claim. He says he does not mean exactly a 70% chance of human extinction. He means roughly a 70% chance that the process goes horribly wrong. His failure category can include:

  • human extinction;
  • permanent loss of meaningful human control;
  • AI systems taking power while some humans remain alive;
  • an authoritarian concentration of AI power in a company, state, or small group;
  • war or destabilization caused by the race itself.

This is still an extreme estimate. There is no historical frequency table for superintelligence from which anyone can calculate a reliable 70%. It is a structured personal judgment assembled from beliefs about timelines, alignment, racing incentives, security, and the consequences of losing control. A different judgment at any layer changes the total sharply.

Kokotajlo is not alone in assigning non-trivial probability to catastrophic outcomes, but his number is not a consensus figure. A published survey of 2,778 authors at top AI venues found deep disagreement. Between 38% and 51% of respondents assigned at least a 10% chance to outcomes as bad as human extinction, depending on question wording. At the same time, 68.3% thought good outcomes from superhuman AI were more likely than bad.

Best interpretation: the field contains substantial concern and substantial optimism at the same time. Surveys show beliefs, not proof. Kokotajlo's 70% estimate should trigger examination of assumptions, not automatic acceptance or ridicule.

AGI and Superintelligence Are Not the Same Claim

TermPractical meaning in this discussionWhy it matters
Current frontier AIBroadly capable but uneven systems that still fail, hallucinate, and need supervision.Evidence about today's systems constrains the forecast but does not settle future behavior.
AGIA contested label for systems matching or exceeding humans across a broad range of valuable cognitive tasks.Different definitions produce different arrival dates and different headlines.
SuperintelligenceSystems better than the best humans across relevant domains, potentially running many copies at machine speed.The risk argument depends on a large capability and power gap, not merely a better chatbot.
Automated AI R&DAI performs most of the work required to improve AI systems.This is the feedback mechanism that makes Kokotajlo's short timeline possible.

The interview also compares neural networks to brains. The analogy can help explain why learned systems are harder to inspect than conventional programs, but it should not be overextended. Artificial neural networks are not established replicas of human brains, and surface similarities do not prove consciousness, human-like motives, or a specific future trajectory.

Jobs, Income, Power, and Purpose

Kokotajlo argues that if superintelligence arrives as defined, almost every job becomes technically automatable. That conclusion is conditional: if a system is better, faster, and cheaper than the best humans at every relevant task, labor demand changes radically by definition.

Technical capability is not the same as instantaneous economic replacement. Adoption can be slowed by regulation, liability, trust, integration costs, physical capital, customer preference, unions, security, and organizational inertia. Some roles may remain human by law or by social choice even when automation is possible.

The interview correctly pushes the discussion beyond "which job survives?" to four harder questions:

  1. Ownership: who owns the models, chips, robots, land, and energy producing the new wealth?
  2. Distribution: how does income reach people if wages stop being the main route?
  3. Power: can a few labs or states convert an intelligence lead into durable political control?
  4. Purpose: what institutions, relationships, and forms of contribution matter when employment is less central?

AI 2040's Citizen's Dividend is one proposed answer to distribution, not a forecast suitable for personal financial planning. Its current economics supplement models roughly $1 million per US person in 2035 and around $10 million in 2040, in 2025 dollars, under extremely aggressive assumptions about automation, output growth, permits, and redistribution. My separate AI 2040 source audit examines those assumptions in detail.

What Plan A Actually Proposes

AI 2040: Plan A is a positive policy scenario published by the AI Futures Project. Its authors explicitly say it is primarily a recommendation, not their best guess about what will happen. The scenario imagines the United States and China reaching an agreement in 2029, avoiding full automation of AI R&D in 2030, pausing at top-human-expert AI in 2035, and proceeding to superintelligence in 2040 after more safety and governance work.

PrincipleOperational ideaHard question
Buy timeSlow frontier capability work when safety confidence is inadequate.Who defines the threshold, and how is a pause enforced?
Research transparencyMake most frontier AI research visible across approved projects and governments.How do you share enough to verify safety without spreading dangerous capability?
Diffuse AIAllow many companies and countries to remain near the frontier instead of one winner holding all power.Does diffusion reduce domination or increase proliferation and attack surface?
ReversibilityLimit algorithmic progress and preserve the ability to disable frontier compute after defection.Can mutually assured compute destruction remain credible without causing escalation?

The proposal includes declarations of major compute holdings, supply-chain audits, reciprocal inspections, inference-only configurations, tighter control of frontier training, and much greater government technical capacity. It is ambitious because the coordination problem is ambitious. A plan that works only when rivals trust one another would be inadequate; Plan A therefore tries to make compliance observable.

The Strongest Objections

  1. The capability timeline may be too short. Long-horizon reliability, robotics, energy, chip supply, scientific experimentation, and organizational deployment may slow the feedback loop.
  2. Forecasting from trend curves can miss regime changes. Scaling laws, revenue growth, and benchmark progress need not continue smoothly.
  3. Risk probabilities are highly model-dependent. The 70% number combines many uncertain conditional probabilities that are difficult to calibrate.
  4. Alignment is not the only possible bottleneck. Security, institutions, economic concentration, military use, misinformation, and ordinary software failures can matter before superintelligence.
  5. Plan A may be politically unrealistic. US-China inspection, deep research transparency, globally distributed frontier projects, and credible compute deterrence require exceptional state capacity and cooperation.
  6. Transparency can create proliferation risk. Publishing frontier research may help verification while also helping irresponsible or covert actors.
  7. The economic abundance model is unusually aggressive. Full automation does not automatically produce fair distribution, stable governance, affordable land, or personal freedom.

None of these objections proves that catastrophic risk is negligible. They identify where disagreement lives. A serious critic should say which link breaks, by how much, and what governance follows from the remaining risk.

How to Read an AI Forecast Without Becoming Gullible or Numb

  1. Label the statement. Is it an observed fact, personal probability, scenario condition, model output, or policy recommendation?
  2. Write the causal chain. Do not debate "AI doom" as one blob. List the steps from current capability to the claimed outcome.
  3. Find the load-bearing assumptions. In this case: automated AI R&D, fast feedback, weak alignment, race incentives, and inadequate intervention.
  4. Ask what would change the forecast. Slower task-horizon progress, poor research automation, hard compute constraints, reliable control methods, or enforceable governance should move the probability.
  5. Track milestones, not vibes. Measure autonomous research cycles, error recovery, time horizons, security incidents, compute concentration, and policy capacity.
  6. Compare with alternatives. Read optimistic, skeptical, and middle-path forecasts, and note where they use different definitions.
  7. Update publicly. A good forecaster records dates, probabilities, misses, and revisions. AI 2027's changelog and clarification are useful examples.
Quarterly AI forecast check

Claim: ______________________________________
Current probability: ________________________
Evidence that increased it: _________________
Evidence that decreased it: _________________
Key missing evidence: _______________________
Next observable milestone: _________________
Action justified now: _______________________
Action not yet justified: ___________________

What a Non-Researcher Can Do Now

  • Learn the systems you already use. Understand their limits, permissions, data flows, costs, and failure modes.
  • Keep consequential actions reviewable. Require approval for deployments, payments, hiring, medical decisions, legal filings, and destructive operations.
  • Build portable skills. Domain judgment, verification, communication, relationship-building, and system management remain useful across many timelines.
  • Avoid single-point dependence. Keep exports, backups, provider alternatives, and human fallback procedures.
  • Ask institutions concrete questions. What evaluations, incident reporting, whistleblower protections, security standards, and deployment thresholds exist?
  • Ask political candidates about AI governance. Move beyond slogans to compute oversight, labor transition, competition, liability, and international coordination.
  • Do not build your life around one forecast. Prepare for faster change without abandoning education, family, work, or long-term plans on the basis of a single scenario.

Kokotajlo's own closing point is not that action is futile. He says that if he believed it were already too late, he would stop this work and spend time with his family. His reason for publishing scenarios is that decisions still matter.

Video Chapters

  1. 00:00:00 - Introduction
  2. 00:02:34 - Why this AI mission could affect everyone
  3. 00:04:21 - Why the average person should care
  4. 00:08:29 - Are AI experts exaggerating?
  5. 00:10:06 - Why Daniel joined OpenAI and what he saw
  6. 00:13:24 - Why he left OpenAI
  7. 00:15:53 - Inside OpenAI during the ChatGPT launch
  8. 00:17:16 - The $2 million non-disparagement controversy
  9. 00:19:30 - Is full AI automation arriving faster than expected?
  10. 00:24:19 - The AI 2027 forecast
  11. 00:26:33 - AGI versus superintelligence
  12. 00:27:17 - Robots in everyday life
  13. 00:30:23 - AI and the human-brain analogy
  14. 00:36:14 - Can AI be creative?
  15. 00:41:15 - The probability of catastrophic outcomes
  16. 00:47:58 - What AI could do to jobs
  17. 00:54:45 - Skills that may still matter
  18. 01:00:21 - AI 2040 and a possible future
  19. 01:04:36 - The different futures AI could create
  20. 01:09:33 - Could AI become Earth's dominant species?
  21. 01:13:18 - How AI CEOs make decisions
  22. 01:15:37 - AI and the 2028 election
  23. 01:18:58 - Is there a safe way to accelerate AI?
  24. 01:22:11 - The scenario where AI supplies 20% of cognitive labor
  25. 01:25:15 - Should everyone receive an AI dividend?
  26. 01:32:41 - What living through the AI transition could feel like
  27. 01:34:11 - Purpose after job displacement
  28. 01:46:12 - Would he shut AI down forever?
  29. 01:49:56 - What people can do about AI
  30. 01:57:25 - Is it already too late?

Bottom Line

Kokotajlo's interview deserves attention because it connects technical capability forecasts to incentives, institutional power, labor, and international security. It also deserves precision. His 70% estimate is a personal probability of a broad catastrophic outcome, not a measured extinction rate. His 2029 date is a median, not a promise. AI 2027 is a scenario, and Plan A is a recommendation.

The right response is neither panic nor reflexive dismissal. Read the causal chain, inspect the assumptions, track the milestones, demand better evidence and governance, and keep updating. High uncertainty is not a reason to do nothing when the downside could be enormous. It is a reason to be unusually clear about what we know, what we believe, and what would change our minds.

Sources

Common questions

Does Daniel Kokotajlo say there is a 70% chance AI will cause human extinction?
Not exactly. In the interview he clarifies that roughly 70% is his probability that advanced AI goes horribly wrong. Human extinction is one possible outcome inside that broader category, alongside loss of human control, an AI takeover, or another severe catastrophe.
Did Daniel Kokotajlo give up $2 million to leave OpenAI?
He refused to sign exit paperwork containing a non-disparagement clause while believing that roughly $2 million in vested equity was at risk. OpenAI later changed its policy, and Kokotajlo has said that he ultimately kept the equity. The accurate description is that he was prepared to lose it, not that he permanently lost it.
When does Daniel Kokotajlo expect superintelligence?
In this interview he gives 2029 as his current median, meaning a 50% probability by then. He also says it could take significantly longer. This is his personal forecasting judgment, not an industry consensus or a scheduled event.
Is AI 2027 a prediction that will definitely happen?
No. The authors call it a concrete scenario representing their best guess at publication, with a race ending and a slowdown ending. The site explicitly says the year 2027 was a modal year, while author medians were longer, and it publishes updates and corrections.
What is AI 2040 Plan A?
Plan A is a policy recommendation expressed through a scenario. It proposes an international agreement to slow frontier AI development, verify major compute holdings, make advanced AI research unusually transparent, diffuse capability across more institutions, and preserve the ability to reverse course before superintelligence.
Do most AI researchers agree with Kokotajlo's 70% estimate?
No single probability represents the field. A survey of 2,778 researchers found wide disagreement: many assigned meaningful probability to extinction-level or similarly severe outcomes, while most still expected good outcomes from superhuman AI to be more likely than bad. These are expert judgments under deep uncertainty, not measured frequencies.
Should people stop using AI because of this interview?
Kokotajlo does not argue for an individual boycott. A reasonable response is to build AI literacy, track concrete capability milestones, preserve human approval for consequential work, ask employers and political candidates about governance, and avoid treating either optimistic or catastrophic scenarios as settled fact.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call