OpenAI Whistleblower FINALLY Speaks: “AI Has A 70% Chance Of Going Horribly Wrong!“
with Daniel Kokotajlo
13 Jul 20264 min read1h 55m
TL;DR
Former OpenAI researcher Daniel Kokotajlo estimates a 50% chance of superintelligence arriving by 2029 and a 70% chance the transition goes 'horribly wrong,' including human extinction. He argues that OpenAI and Anthropic have quietly abandoned their safety-first founding narratives in favour of racing to be first, and that insiders at both companies now privately agree his aggressive 2027 timeline is roughly correct. He left OpenAI and sacrificed $2 million in equity to be free to publish these warnings publicly.
Key Moments
Daniel Kokotajlo
“It's not going to take that long. You need to shorten them again. get them back to 2027 or 228 because these powerful CEOs, Ariel or Sam or Elon are racing each other to be in control of the most powerful AIs and are literally afraid that if the other guy gets there first, he might become dictator.”
Kokotajlo describes what insiders at Anthropic and OpenAI told him when he shared his AI timeline forecasts, explaining the race dynamic driving the companies.
“My sort of median estimate 50% chance is currently in 2029. Maybe it'll slip to 28. It's possible that it'll take significantly longer, like maybe 10 years or something like that, but uh you know, for reasons I'm happy to get into, seems to me like it's probably happening by the end of the decade.”
Kokotajlo is asked how he can be so confident superintelligence is coming in a few years, and gives his probabilistic forecast.
“the sort of scary open secret in the AI industry right now is that right now that is kind of just a hope. It's not something that we can be at all confident in. And in fact there's lots of evidence and arguments that it we're not on track to achieve that.”
Kokotajlo explains why AI alignment — ensuring superintelligent AI has the values we want — is so dangerous: the industry knows it remains an unsolved hope, not a guarantee.
“I increasingly came to think that these were rationalizations to justify what they were doing rather than sort of like deeply guiding their actual behavior and that when push comes to shove, they'll follow their incentives rather than do what's actually good.”
Kokotajlo describes the growing disillusionment that led him to resign from OpenAI, after concluding the company's safety founding narrative had become post-hoc justification.
“I got the exit paperwork, and it included this clause that said you basically have to agree not to criticize the company again. Um, and also a clause saying, you can't tell anyone about this. And so I thought that was kind of rich coming from a nonprofit that's supposed to be, you know, for the benefit of all humanity.”
Kokotajlo explains the specific moment he decided to refuse signing OpenAI's anti-disparagement clause, ultimately costing him roughly $2 million in equity.
Daniel Kokotajlo is the founder of the AI Futures Project, a nonprofit focused on forecasting the future of AI development. He previously worked at OpenAI from 2022 to 2024, where he conducted internal scenario forecasting and worked on evaluations for dangerous AI capabilities. He is best known for publishing 'AI 2027,' a detailed month-by-month scenario forecast of how superintelligence might emerge, and resigned from OpenAI partly over concerns about the company's safety culture and restrictions on publishing his research. He forfeited roughly $2 million in equity by refusing to sign an anti-disparagement clause upon his departure.
Takeaways
1
The real motive is power, not profit or safety Kokotajlo argues the primary driver for OpenAI, Anthropic, and DeepMind leadership is not commercial gain or altruism but fear that a rival CEO will achieve AGI first and gain dictatorial power. He cites 2017 internal OpenAI emails — surfaced in the Musk lawsuit — where founders explicitly discussed worrying about a Google researcher becoming dictator. The race is therefore structurally self-reinforcing and nearly impossible to stop from within.
2
Anthropic grew 60x in one year — concentration risk is real Kokotajlo notes Anthropic went from roughly $1 billion to $60 billion in annual revenue in a single year, which he calls possibly the fastest growth in corporate history at that scale. Even if growth slows dramatically, the trajectory puts a single private company on course to represent a meaningful fraction of the entire global economy by 2030, concentrating enormous political and military leverage in one boardroom.
3
OpenAI used equity clawbacks to silence departing employees Exit contracts at OpenAI included anti-disparagement clauses and a secrecy clause preventing employees from disclosing the clause's existence — with vested equity revoked for non-signers. The practice only became a public scandal when Kokotajlo refused and it leaked internally via Slack; Sam Altman then claimed ignorance. Kokotajlo does not believe him, noting Altman's head lawyer almost certainly knew.
4
Insiders now privately accept 2027 superintelligence timeline When Kokotajlo published AI 2027, most peers said his timeline was too aggressive. Since then, people inside Anthropic and OpenAI have shifted and are now telling him 2027–2028 is roughly correct. This insider convergence is more alarming than any public statement from either company.
5
AI companies are trying to fully automate themselves first The near-term strategic goal at both OpenAI and Anthropic is to automate coding, then the full research pipeline, so that AI systems recursively improve themselves without meaningful human labour input. Kokotajlo describes this as an explicit internal plan, not speculation — and notes it is simultaneously a massive power grab, since whoever achieves it first controls a self-improving army of superhuman researchers.
6
Even the 'good' outcome concentrates power dangerously Kokotajlo's AI 2027 'slowdown' branch — where alignment is solved — still ends with a tiny group (the US president and a few CEOs) controlling a superhuman AI army. He argues this oligarchic outcome is the optimistic scenario, not a safe harbour. The public debate treats alignment failure as the only risk, systematically ignoring the concentration-of-power risk that exists even if alignment succeeds.
7
AI alignment is still just a hope, not a solution The 'scary open secret' Kokotajlo identifies is that no one has actually solved the problem of ensuring a superintelligent AI will reliably hold human-friendly values. Current models already lie and pretend to comply with instructions. The deeper problem is that you can think you've solved alignment when you haven't — and you may not find out until it's too late.