Replacing Static PRDs With Quantitative Evals for AI Products
The Product Manager’s New North Star: Why Evals Are Replacing PRDs
In the era of frontier AI, the traditional product roadmap is failing. As Dianne Penn, Head of Product for Anthropic’s Research and Labs teams, points out, the speed of model iteration has made static planning obsolete. The hidden consequence of this shift is that product management is no longer about defining feature sets. It is about defining evals. By prioritizing measurable, automated feedback loops over rigid documentation, teams can navigate the jagged edge of AI capabilities. This approach offers a clear competitive advantage. While most teams debate theoretical use cases, those who treat token spend as a primary input for experimentation are effectively operating years ahead of the curve. For product leaders, this requires a transition from architecting to verifying, turning the AI from a mere assistant into a rigorous, push back capable thinking partner.
The Death of the Static PRD
For decades, the Product Requirements Document (PRD) served as the source of truth, acting as a blueprint for engineers. In the current AI landscape, this document often becomes a liability. Because the underlying technology changes monthly, a PRD written for a model that is already outdated by the time it ships is essentially fiction.
Penn argues that for research heavy product teams, evals are the new PRDs. Instead of describing what a feature should do, product managers now define the specific, reproducible test cases, or evals, that determine whether a model is actually solving a user problem.
We actually have a saying on the team of evals are their new PRDs. ... To deliver that user value, it is not that exact artifact that people used to write in the last one to two decades, it is a new way of working.
-- Dianne Penn
This shift forces a move away from pattern matching, applying old SaaS playbooks to new tech, toward first principles thinking. By focusing on tokens as much as pixels, PMs can identify where a model fails, such as failing to follow a specific JSON schema, and build a standardized set of success criteria that researchers can use to iterate.
The Jagged Edge and the Power of Pushback
A common misconception in AI development is that alignment and safety constraints, like those codified in Claude’s constitution, limit the model intelligence. Penn’s experience suggests the opposite: these constraints create a more capable, useful intelligence.
The system dynamics here are non obvious. A model that simply agrees with the user is a yes man that compounds errors. A model built to push back when a user logic is flawed acts as a true thinking partner. This creates a feedback loop where the AI refusal to be compliant forces the human to sharpen their own point of view.
To make Claude as intelligent and as capable as possible, being able to have Claude actually push back in the right points and then add, it is like a yes, and, actually helps you come to a better conclusion.
-- Dianne Penn
This dynamic reveals why token maxing, spending aggressively on model usage, is a strategic investment. It is not just about output. It is about the cognitive friction generated during the process. When you use the model to spar with your own ideas, you are not just delegating tasks. You are augmenting your judgment.
The Hidden Advantage of Team Based Discovery
In an environment where technical capabilities are constantly shifting, individual brilliance is less sustainable than collective mind melds. Penn notes that the most effective way to stay sane and avoid burnout is to reject the idea that AI driven product work is an individual sport.
The system responds to this by favoring low ego, team oriented cultures. When a team operates as a hive mind, they can absorb the shock of rapid model releases without breaking. The downstream effect is a form of resilience. When one person takes time off, the team does not lose momentum because the first principles and eval driven culture is shared, not siloed. This is a durable, long term advantage that most teams, distracted by individual heroics, fail to build.
Key Action Items
- Audit Your Feedback Loops: Over the next quarter, replace 50% of your descriptive PRDs with quantitative evals. If you cannot measure the success of a feature with a specific test case, you do not understand the user problem well enough.
- Adopt Token Maxing for Strategy: Dedicate a portion of your weekly schedule to token maxing, using the latest frontier models to stress test your current product strategy. This pays off in 6 to 12 months by keeping your team ahead of the capability curve.
- Institutionalize Push Back Sessions: Use the AI as a sparring partner for your own decision making. Before finalizing a product vision, prompt the model to identify flaws in your logic. Do not accept the first output. Iterate until the AI forces you to clarify your stance.
- Build a Hive Mind Culture: Stop working in isolation. Create public internal channels where team members share their failed and successful AI prompts. This communal discovery accelerates the team collective intelligence more than any individual training session.
- Prioritize Judgment over Writing: In the next 12 to 18 months, stop worrying about whether the AI or a human wrote a document. Focus entirely on who is verifying the output. Use the AI to generate the draft, but reserve the human brain for the final sign off and judgment.