Building Proprietary Evaluation Frameworks to Ensure AI ROI

Original Title: Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan

The Evaluation Gap: Why AI Capabilities Are Outpacing Our Ability to Measure Them

Rayan Krishnan, Ben Horowitz, and the a16z team discuss the systemic failure of current AI benchmarking. The core issue is that public benchmarks have become gimmicks that encourage models to hack tests rather than show real utility. This creates a dangerous information gap: labs report high scores while businesses cannot quantify the actual return on their AI investment. The result is a misallocation of capital where token costs often exceed human labor costs, yet performance remains unclear. For leaders, the advantage lies in building proprietary, task-specific evaluation frameworks rather than following public leaderboards. Those who treat evaluations as a core business competency will be the only ones capable of moving from theoretical potential to operational reality.

The Illusion of Performance and the Pay-to-Pass Trap

The industry is caught in a loop where model labs and benchmark providers share the same incentives. Because public benchmarks are open, they are easily gamed. Krishnan notes that a model might show incredible results on public tests while failing on private, high-signal benchmarks. This disconnect creates a pay-to-pass environment similar to the audit failures of the Enron era, where the entities responsible for evaluating performance are also incentivized to maximize the model's perceived success.

There is a huge disconnect between what was self-reported based on these open benchmarks and then what we were actually finding with our higher quality, higher signal benchmarks.

-- Rayan Krishnan

When companies rely on these public metrics, they are building their AI strategy on a foundation of gimmick data. Over time, this shifts the incentive structure away from solving real-world problems and toward optimizing for the specific quirks of a public test.

The Existential Risk of Misvalued Intelligence

The most significant insight from the discussion is that the AI ROI problem is already a threat to large enterprises. Krishnan describes a Fortune 10 company where engineers are limited by arbitrary token budgets, leading to dead periods in the afternoon when teams stop working because their rate limits have been hit.

The system responds to these constraints with inefficiency: teams spend 10 times more on tokens than on the salaries of the employees using them, yet they lack the framework to know if that spend is generating value.

We are in this world where it is still very unclear what ROI looks like and how to value this intelligence that is being used. And so as you talk about the existential concern for enterprises, I think it is this kind of direction we are shifting in where token spend may start to eclipse salary spend.

-- Rayan Krishnan

This creates a competitive chasm. Companies that fail to make their evaluations legible will continue to lose capital on inefficient token usage. Conversely, those that build internal, repository-specific benchmarks can identify which models are optimal for their specific codebase, effectively buying their way out of the messy middle of expensive, generic model usage.

The Always a Higher Peak Systems Dynamic

Krishnan frames AI evaluation as an infinite game. As soon as a benchmark is created, model labs begin hill climbing to beat it. This requires a constant cycle of deprecating old benchmarks and constructing new ones. This is not just a technical chore; it is a requirement for maintaining a competitive edge.

The system is evolving toward agentic workflows, which are models that operate over hours, days, or weeks. Evaluating these is a different challenge than the one-to-one mapping of early image classification. It requires simulating entire cloud environments and infrastructure, not just testing code snippets. The firms that treat evaluation as a dynamic, evolving capability will be the ones that can audit their own recursive self-improvement, while others will be left guessing whether their models are actually advancing or simply hallucinating progress.

Key Action Items

  • Audit your token spend (Immediate): Map your current token consumption against specific project outcomes. If your token spend is approaching salary-level costs, you are likely over-spending on generic frontier models where smaller, task-specific models would suffice.
  • Build proprietary benchmarks (Next 30-60 days): Stop relying on public leaderboards to choose your vendor. Use your own internal GitHub repositories or project workflows to create a private benchmark that reflects your actual business requirements.
  • Implement Task-Based Routing (Next Quarter): Move away from a one-size-fits-all model approach. Use your evaluation framework to route simple tasks to cost-efficient models and reserve frontier models only for tasks where the intelligence delta is proven to be worth the cost.
  • Formalize the Eval function (6-12 months): Treat evaluation as a core engineering discipline. Just as you have QA for software, you need a dedicated, neutral team responsible for evaluating model performance against business-specific rubrics.
  • Prepare for Recursive monitoring (12-18 months): As models become more agentic and capable of self-improvement, start investing in infrastructure that can monitor long-running, asynchronous agent trajectories rather than just single-prompt responses.

---
Handpicked links, AI-assisted summaries. Human judgment, machine efficiency.
This content is a personally curated review and synopsis derived from the original podcast episode.