Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne Penn
with Diane Penn
26 Jul 20265 min read1h 10m
TL;DR
Diane Penn, Anthropic's first technical PM, argues that Opus 3's coding focus in early 2024 was the hidden inflection point that differentiated Claude — and that Opus 45 only succeeded because Claude Code existed as a 'vehicle' for the model's intelligence. She frames evals as 'the new PRDs' and says the best PMs right now should be asking: if Claude 8 comes out, what changes in what users do, and is what I'm building today forward-compatible with that?
Key Moments
Diane Penn
“Evals are the new PRDs.”
Diane is describing how the product role is changing at Anthropic, where evaluation frameworks have replaced traditional product requirements documents as the core driver of user value.
“Opus 45 wouldn't have had that moment without a product like Cloud Code and Cloud Code I think wouldn't have had that type of adoption accelerated without Opus45.”
Diane is explaining why Opus 45 was a major inflection — not just the model itself, but the combination of model and product experience.
“Unless you have the eval unless you have the systems to test um these jumps might actually happen and you don't know.”
Diane is explaining how emerging capabilities in AI models can appear discontinuously — and why evals are essential to catch capabilities you didn't know the model had developed.
Diane Penn is Head of Product for the AI Research and Labs teams at Anthropic, where she joined as the company's first technical product manager over three years ago when the product team had just five engineers. She has helped ship every model Anthropic has released, from Claude 2 through Fable, and has incubated and launched major products including Claude Code, MCP, Skills, Claude Design, and core capabilities like computer use, tool use, and reasoning. Her unique position at the intersection of research and product gives her a front-row seat to how frontier AI models are trained, evaluated, and brought to users.
Takeaways
1
Evals are the new PRDs for AI PMs At Anthropic, the team has replaced traditional product requirements documents with evaluation frameworks as the primary way to drive user value and guide model improvement. If you're building AI products, your ability to write precise, actionable evals is now more important than writing feature specs.
2
Ask 'what does Claude 8 change?' before shipping Diane's standing question to her team is: if Claude 8 arrives, what changes in what users do — and is what we're building today forward-compatible with that future? This forces product decisions to be durable across model generations rather than optimized for today's capabilities.
3
Model plus product vehicle — both required Opus 45 only became a breakout moment because Claude Code existed as a 'vehicle' to deliver its intelligence to users. Diane's framing: 'you need frontier products in order to have frontier models' resonate. A powerful model without the right product experience won't achieve adoption, and vice versa.
4
Coding focus was Claude's hidden early differentiator In 2023, nobody associated Anthropic or Claude with coding. Diane identified that users were starting to write long-form code — not just autocomplete — and made a relatively small training change to Opus 3 that became a major competitive differentiator and attracted Claude's earliest developer enthusiasts.
5
Emerging AI capabilities appear discontinuously — build evals to catch them Scaling law papers show smooth loss curves, but capability graphs show sudden discontinuous jumps — the model goes from being unable to calculate 1+1 to doing it reliably in one leap. Without evals designed to test for these, teams can ship a model that can do something they're completely unaware of, which matters both for product opportunity and safety.
6
Labs bets: strong thesis, loose prototype Anthropic Labs operates with 'strongly held opinions about the theme or area and more weakly held about the exact prototype.' This lets small teams — sometimes starting with a single engineer — pursue discontinuous bets without being slowed by large team coordination, and revisit failed prototypes one or two model generations later when the model may have caught up.
7
Token spend reframed as experimentation output Rather than endorsing Gary Tan's '$100k/year token maxing' framing directly, Diane reframes it: token spend is the input, and experimentation is the real output to optimize for. The most creative thinkers at Anthropic spend heavy time with every new research model version — there is no substitute for direct contact with the technology when it's moving this fast.