Article — Position paper · ○ Open access

Scarcity counts the order, not the consumption

When capital, architecture and value markets decouple, resource allocation follows whichever market still emits a price. Enterprise AI is the current instance.

Jérôme Vetillard · · Twingital Institute · 7 pages · 6 min read
🇫🇷 Lire en français ↓ Download PDF

Enterprise AI in mid-2026 allocates capital, silicon and engineering effort against a signal that no longer tracks realized value. Three markets that should stay separate have come apart: a capital market that prices future capacity, an architecture market that prices what gets built, and a value market that would price outcomes if it could measure them. When the third market emits no price, the first governs the second by default. Compressed to its shortest form, the mechanism reads: scarcity counts the order placed against a plan, not the workload actually served. This is a mechanism, not a bubble call, and its domain of validity is narrow: enterprise deployment under usage-based billing, on the realized figures of mid-2026, strongest where the purchase is a committee decision and the signer is insulated from the meter, weakest where a single owner prices its own cash flow.

Three markets that should stay separate

The deployability thesis held that the hard problem of enterprise AI is institutional insertion, not the model. This piece names the mechanism underneath it. Two definitions fix the argument before they drift. Rent, here, is the operating cost incurred above what the task requires to reach the same useful decision: the surplus, and only the surplus, never the agent and never the complexity a problem genuinely demands. And disproportionately is deliberate: the architecture market answers to talent, regulation, open standards and cloud economics as well as to capital. Capital is simply the loudest voice today, not the only one. The claim generalizes beyond AI, to any setting where value creation is deferred or hard to measure while capital is neither.

Why the shortage measures the order, not the served workload

Two facts sit in the same market. High-bandwidth memory and packaging are effectively sold out, GPU lead times run in quarters, and power has replaced silicon as the binding constraint, with the Apollo and JPMorgan desks describing a supply chain tight through 2026. Against that, Cast AI’s 2026 review of roughly twenty-three thousand enterprise Kubernetes clusters found average GPU utilization near five percent. Read at its scope, and only at its scope, it proves one thing: a large population of enterprise clusters provisioned during the scramble sits idle while the provider bills for it whether it turns or not. The paradox dissolves once one asks what the scarcity is measured against. Memory is sold out against orders, and orders are placed against build targets.

The strongest counter deserves full weight: fiber in the late 1990s and cloud in the late 2000s ran at abysmal initial utilization and preceded their applications by years, so overcapacity may be the normal lag of a latent-demand transition. The disanalogy is the depreciation clock. Fiber was cheap to hold idle; GPU capacity, written down in eighteen to thirty-six months, is not. Sequoia’s David Cahn puts the annual gap between spend and ecosystem revenue near six hundred billion dollars and widening, and it is the gap, not the utilization figure, that is the number to watch. Nvidia sets the clock: Hopper, then Blackwell, then the Rubin generation shown at GTC 2026, each architecture compressing the economic life of the one before it and forcing the owner to fill capacity before it expires.

Rent is the fifth source of weight, and the only one nobody can see

Weight has five sources, and four are legitimate. It can be a property of the problem: identity, security, observability and orchestration are hard, and a serious deployment is heavier than a demonstration. It can be a property of regulation: under the AI Act or the European Health Data Space, an architecture that logs, isolates and keeps a human in the loop is legal optimality, rational on an axis that is not the economic one. It can be optionality, the entry ticket that stabilizes internal skills and restructures enterprise data, a cost that is tuition rather than rent, and only the buyer knows which it bought. It can be immaturity, an agent stacked on an agent under deadline. And it can be rent, the surplus as defined. The difficulty is that the fifth is indistinguishable from the other four without a baseline almost no one runs, the cheaper architecture set alongside to see whether it would have sufficed.

One illustration fixes the point without carrying the argument. A calibrated toxicity model mapping a molecule to fourteen endpoints, of the kind built as ToxTwin on the Institute’s own terrain, predicts, reports a confidence and stops, at a fraction of what a generic agent would spend to reach the same output. It is one case where the baseline is obvious. In most enterprise settings it is not, and that, not the existence of the frugal option, is the problem. The thesis does not claim enterprise architectures are too heavy. It claims that nothing in the buyer’s instruments separates the weight that is warranted from the weight that is rent.

The deployer’s incentive points one way

The architecture reaches the client through two utility functions that happen to align. The hyperscaler’s professional-services arm inherits the owner’s imperative, utilization. The system integrator runs on billable hours, maintenance, support and change management, none of which reward the lighter architecture, and it is often the integrator, closer to the client’s estate, who is the more decisive of the two. Neither requires bad faith: the bias is a property of the incentive field, not of anyone’s character. Set the deployed narrative against the record. MIT’s NANDA initiative, in The GenAI Divide (2025), found about ninety-five percent of generative-AI pilots with no measurable P&L impact and about five percent reaching rapid revenue acceleration; Boston Consulting Group’s Widening AI Value Gap (September 2025) put sixty percent of firms at no material value; McKinsey’s late-2025 survey found enterprise EBIT impact at fewer than four in ten adopters. The cost the narrative omits is institutional: evaluation, observability, governance and the human review a confidently wrong system forces. An architecture can be cheap in tokens and ruinous in organization, and a customer-ROI figure that counts neither is not measuring a return.

The missing ruler: cost per useful decision

Akerlof’s Market for Lemons (1970) showed that when buyers cannot observe quality the market selects for what is observable, and unseen quality cannot be paid for. Here the unobservable is the line between productive consumption and rent; the observable is the demonstration and the benchmark, which is precisely the weight the supply side is rewarded to show. The ruler is genuinely hard to build, for four reasons that compound: causality is diffuse, temporality is long, attribution is near-impossible when a human corrects the model, and the externalities land on morale, retention and speed, none of which finance books. Name the instrument, and it is a family rather than a single number: cost per useful decision, a ratio whose denominator counts decisions whose downstream value exceeds the cost of verifying them, whose numerator sums the total lifetime cost of producing them, compute plus oversight plus governance plus verification plus operations. The denominator is domain-specific and cannot be otherwise. Its sharpest limit is that whoever defines the useful decision controls the ratio: if the integrator captures that definition, the metric becomes marketing rather than discipline, and the asymmetry simply moves from the token to the word useful.

What the decision-maker does now

This argument proposes a mechanism and illustrates it; it does not measure it. Akerlof, Solow, Cast AI, MIT and Sequoia illustrate a chain that has not been tested as a chain, and the honest status of the claim is a hypothesis with a specified test: measure cost per useful decision across matched deployments with and without the metric, and read the difference. Solow’s 1987 condition, the computer age visible everywhere except in the productivity statistics, applies again with one novelty. Under a license the unmeasured period cost was fixed; under usage-based billing the meter runs while you wait. The consequence for the decision-maker is neither adoption on faith nor withdrawal in panic, the same error facing opposite directions. It is to build the ruler faster than the market standardizes it, and to refuse the two substitutions the current phase runs on: the shortage against a plan read as a shortage against a need, and committed spend read as realized return. The organizations that priced their deployments in useful decisions will still be standing when the meter has run its course. The organizations that priced them in conviction will discover what the invoice was for.

A customer ROI that is asserted and never measured is not a return. It is a narrative with an invoice attached.

Full article, with all sources and the complete six-part argument, in the PDF below (7 pages).

Read the document