All-In

Open Source Wins, AGI Is Here, and Scorsese's AI Toolkit with CEOs of Cerebras & Black Forest Labs

with Jason Calacanis, David Sacks, Chamath Palihapitiya, David Friedberg
10 Jul 2026 5 min read 1h 10m

Cerebras CEO Andrew Feldman argues we have already hit AGI by any definition from 20 years ago, and that the next leap — super intelligence — will come from recursive 'loop maxing,' where models iteratively improve their own reasoning. Feldman reveals Cerebras has a $25 billion backlog, and explains that his chips' blisteringly fast inference is what makes extended reasoning runs (25–48 hours) tractable. The conversation also tackles open-source model sovereignty, the Anthropic government standoff over staged model releases, and why every nation from Kazakhstan to Armenia is now building out data centers.

Andrew Feldman
“what we're talking about now are data centers that are in the next several years going to use more power than the previous 50 years on Earth took.”
Feldman is describing the unprecedented physical scale of the current AI infrastructure buildout to Jason Calacanis.
▶ 2:16
Andrew Feldman
“The irony is unlike many sort of exciting times in technology. They're trying to capture yesterday's demand, right? The demand is way outstripping our ability to build data centers and to fill them with hardware.”
Feldman explaining that AI compute customers like OpenAI and Anthropic are so desperate for capacity they order chips before they are finished being made.
▶ 3:38
Andrew Feldman
“And uh my view is in the next 18 months we'll be way over 2x”
Feldman describing Cerebras's internal performance trajectory and why a newer architecture can outpace traditional Moore's Law doubling.
▶ 13:28
Andrew Feldman
“AGI I think I suspect you'll agree with me that we've hit it. we just haven't exactly deployed it fully. We have artificial general intelligence now.”
Feldman making the direct claim that AGI has already been achieved, framed against any definition that would have been used 20 years ago.
▶ 29:27
Andrew Feldman
“powerful recursive gains are exponential, right? you get better, you do it again. And if you continue to get gain, the slope of that curve is so steep.”
Feldman explaining the mathematical logic behind 'loop maxing' and why recursive self-improvement is the path to super intelligence.
▶ 32:16
All-In is a podcast hosted by Silicon Valley veterans Jason Calacanis, David Sacks, Chamath Palihapitiya, and David Friedberg. The hosts debate the biggest stories in tech, venture capital, politics, and markets. This episode features Andrew Feldman, CEO of Cerebras Systems, discussing the AI infrastructure buildout and the state of reasoning models.
1
Fast inference chips make 48-hour reasoning runs viable Extended reasoning runs — where a model deliberates for 25 to 48 hours — produce qualitatively different and vastly better outputs, not just incrementally better ones. Cerebras's speed advantage (cited as ~15x faster) means what would take weeks of inference time on slower hardware becomes practical in a single day. This makes 'token maxing' and 'loop maxing' economically and operationally feasible for real workloads.
2
Open-source AI sovereignty is now a board-level decision Regulated industries — finance, healthcare — are increasingly choosing open-source models deployed on-premises to avoid data leakage and dependency on frontier model providers. Feldman notes that in practice the only meaningful open-source options today are OpenAI's OSS 12B or Chinese models (GLM, Kimi, Qwen), creating a gap that domestic US open-source efforts need to fill. The government standoff over Anthropic's staged model release accelerated this conversation, especially in Europe.
3
Prompt engineering is dying; intent understanding is here Models like OpenAI's Fable and o3 (referred to as '56') are increasingly inferring user intent without requiring precise prompt construction, marking a fundamental shift from earlier generations where small wording changes dramatically altered outputs. Feldman and Calacanis both observe that the AI now proactively suggests better framings, additional charts, or follow-up questions the user didn't think to ask. Adding explicit instructions like 'check your work and tell me what I haven't considered' dramatically improves output quality.