Mastering Token Throughput Efficiency for Sustainable AI Profitability
The $1.4 trillion question facing Big Tech is not whether these companies can build the necessary infrastructure, but whether they can capture the value it creates. While the scale of investment is unprecedented, with compute capacity expected to quadruple by 2028, the real competitive advantage lies in moving from training models to high-efficiency inference. By mapping the return on invested capital (ROIC) across three business models, Brian Nowak shows that AI economics depend less on total capital expenditure and more on token throughput efficiency. For investors and operators, the lesson is clear: the winners will be those who master the unit economics of token processing, rather than those who simply own the largest data centers. This analysis offers a framework to distinguish between AI players building sustainable moats and those simply burning cash.
The efficiency trap: Why throughput defines the bottom line
Conventional wisdom treats massive data center spending as a binary bet on demand. However, system dynamics suggest otherwise. The true bottleneck, and the primary driver of ROIC, is how efficiently infrastructure processes tokens. As Nowak notes, the ability to squeeze more tokens per GPU per second is the variable that dictates long-term profitability.
"This is why continued improvements in chips and software to drive higher token throughput -- or more tokens per GPU per second -- are critical to the long-term unit economics across this AI ecosystem."
-- Brian Nowak
This creates a feedback loop: as token throughput increases, the cost per unit of intelligence drops, which drives higher adoption and revenue. If a firm fails to optimize this throughput, their ROIC will likely compress, regardless of how much capital they deploy. The advantage here is delayed; companies that prioritize software-level optimizations for inference today are building infrastructure that will remain profitable even if compute rental prices fluctuate.
Vertical integration vs. the infrastructure tax
The decision to own or rent infrastructure is a structural choice that alters the risk profile of an AI business. Nowak’s analysis indicates that when an AI lab owns its own infrastructure, it captures a larger share of the value chain, leading to 40%+ ROIC.
When a developer rents infrastructure, they pay an "infrastructure tax" to the cloud provider. While this model still yields a 25% post-tax return, it creates a dependency where the developer’s margins are tied to the pricing power of the infrastructure provider.
"While this lowers their returns on invested capital because another provider takes a piece of the unit economics, our base case still produces roughly a 30 percent incremental operating margin and 25 percent post-tax return potential."
-- Brian Nowak
This suggests that the most durable companies will be those that successfully transition from renting to owning, or those that achieve such high efficiency that the "infrastructure tax" becomes a negligible part of their total cost structure.
The shift from training to inference
The most significant dynamic is the impending transition of infrastructure utility. Currently, much of the $1.4 trillion is flowing into training models. However, the true test of this investment is the shift toward inference, or the actual serving of products to customers.
The industry is currently in a build-out phase, but the payoff depends on the transition to serving. If infrastructure remains locked in a cycle of training, ROIC will remain theoretical. As the industry moves toward inference, revenue streams will become more predictable. Competitive advantage belongs to those who successfully bridge the gap between training massive models and deploying them in high-margin, scalable products.
Key action items
- Audit token throughput: Over the next quarter, prioritize investments in software and chip-level optimizations that increase tokens per GPU per second. This is the primary lever for protecting ROIC.
- Evaluate infrastructure dependency: Assess whether your business model relies on renting compute. If so, calculate the "infrastructure tax" impact on your 18-month margin projections and explore vertical integration strategies.
- Monitor inference scaling: Shift internal metrics from training progress to inference efficiency. This is where long-term, sustainable revenue will materialize.
- Stress-test ROIC scenarios: Do not rely on base-case rental pricing. Run models at the low-20% ROIC range to ensure the business model remains viable even if market prices for compute face downward pressure.
- Prioritize deployment over training: Over the next 12-18 months, ensure that infrastructure spend is increasingly tied to revenue-generating inference products rather than experimental training cycles.