Article — Position paper · ○ Open access

The Token Is Not the Unit

Cost per Useful Decision, or why the falling cost of inference shifts the charge toward proof and risk.

Jérôme Vetillard · · Twingital Institute · 11 pages · 9 min read
🇫🇷 Lire en français ↓ Download PDF

The note that opens this sequence established that an endorsed decision is the first governable object in enterprise AI. It left open the question this one treats: what that decision costs. Vendors answer by the million tokens, and a reductive reading of FinOps applied to AI, dominant today, has adopted that commercial unit and turned it into the primary proxy for economic efficiency. The problem is not that the token is measured; it is that it serves as the measure of a system for which it is only the input, an input bound to become more substitutable and less differentiating than the institutional capabilities that surround the decision. The domain of validity is narrow, and it is stated before the argument: the frame begins when the output of the system materially modifies a right, an obligation, an exposure or a trade-off borne by the institution. Batch reformatting stays outside it. A simple classification enters it as soon as it denies access, prioritises an investigation, excludes a file, steers a diagnosis or raises a premium.

Three units are conflated: billing, monitoring, judging

The ambient discourse stacks three units that nothing obliges to coincide. The unit of billing is the resource purchased from the vendor. The unit of management is what the operator monitors day to day. The economic unit is that by which efficiency is judged. The token legitimately occupies the first two places and usurps the third. Outside the committing domain it may occupy all three without damage, and this note claims nothing there. Inside it, the substitution is expensive: it makes a system be judged on its most commoditised input, at the precise moment that input stops differentiating anything at all. The system is sometimes billed by the token. It is always steered by the decision.

The decision is the minimal unit of commitment, not the ultimate unit of value

A firm does not turn to AI in order to accumulate generated tokens, and no more to produce value directly: it turns to AI to produce a result that commits it. Two objects must be separated, the minimal unit of commitment and the ultimate unit of value creation. The decision is the first, not the second. It is the first institutional state to which the organisation can attach a commitment, a liability and a cost, and a determining algorithmic recommendation suffices, even without the legal form of a decision. Final economic value depends on the later execution of the business process, over which the system no longer has any hold. Conflating the two produces two symmetrical errors: crediting AI with all the downstream value, or denying it any cost of its own upstream. The chain reads compute, inference, decision, action, outcome, value; the perimeter of CUD stops at the third position. The frame covers the full set of regimes of decisional participation, and this generality is deliberate: where human validation is routine or weakly adversarial, it does not neutralise the economic influence of the system, it displaces its formal signature.

Three charges in the numerator, one convention in the denominator

Cost per Useful Decision is written CUD = (O + K + H) / D^U, over one period, one perimeter and one cohort. The numerator gathers the whole of costs into three exclusive aggregates. O is the operating cost of the delegation regime, compute (O_tech) and supervising humans (O_human) taken together. K is the cost of the governance and proof capability, the architecture that maintains the attributability, the integrity and the revisability of what has been decided. H is the economic charge of risk, the financial translation of exposure. No option cost, vendor dependency, reversibility, exit or transition cost escapes the frame: each is imputed to one of the three according to its nature, under a stable allocation convention.

The denominator does not count inferences, it counts useful decisions: institutionally closed, compliant with the quality thresholds, having not exceeded the remediation budget set for their class, and still compliant at the end of an observation window Δ. Thirty days for certain credit decisions, ninety for certain claim settlements, the choice depending on the timing of appeals and the delay in loss materialisation. Three precautions govern any calculation. The numerator mixes distinct cost statuses, observed costs, imputed costs, estimated losses, capital charges, and confusing them means taking a management construct for a measure. Numerator and denominator must share their horizon, failing which the ratio measures nothing. And minimum quality is a conjunction of non-compensatory thresholds: a system does not redeem defective traceability with better latency.

What would refute the thesis, and what does not

Two statements must be distinguished, and conflating them would make the proposition either trivial or untestable. The first is analytical: with the rest of the cost structure held fixed, reducing the technical component O_tech alone mechanically raises the share of non-technical charges. This is an identity; it proves nothing. The second is empirical, and it is the thesis: in real deployments, a fall in O_tech does not produce a proportional fall in the absolute level of O_human + K + H. These charges are governed by regulation and exposure, not by the price of the GPU, and their elasticity to the cost of compute is low. The refutation condition is clear: if, at constant perimeter, volume, quality and liability regime, that absolute level fell at the pace of O_tech, the thesis would be false. This note proposes the measurement frame and the test condition; it does not establish the low elasticity empirically.

The figures that follow are illustrative, in normalised units, and verify a mechanism without proving a general property. One class of credit granting decisions, two regimes. Assisted architecture, the AI recommends and the human decides: compute 1, supervision 5, proof 2, risk 2, giving 10. Automated architecture, the AI decides and the human handles the exception: compute 2, residual supervision 2, proof 4, risk 3, giving 11. First observation, counter-intuitive: before any commoditisation, automation costs more per unit. Now let compute tend toward free: the assisted regime moves to 9, the automated one to 9. The equality holds because the figures were set that way; it is not a consequence of CUD. What it does show holds: automation lodges the bulk of its cost in proof and risk, pockets that no progress on inference optimises. A warning completes the thesis: a fall in the total cost does not refute it, since that total may decline through mutualisation of the proof base, through contractual transfer of risk, or through a simple rise in volume. The thesis bears on composition and on the residual, not on the level.

Where Shannon had a unit, CUD has a convention

The founding gesture is a separation. Just as information theory separated the quantity of information from the physical medium that carries it (Shannon, A Mathematical Theory of Communication, 1948), the economics of AI must separate the decision produced from the compute resource that makes it possible. The analogy is useful and asymmetric, and the asymmetry is more interesting than the resemblance. Shannon’s quantity of information could be formalised independently of its medium, in a natural unit, the bit, with no institutional convention. CUD has no such support. The useful decision is not a natural magnitude: it is defined by a conjunction of non-compensatory quality thresholds, a maturation window and a remediation budget, all set by governance. Its robustness therefore does not derive from a unit, but from the stability of its definitions, its thresholds and its imputation rules. This is not an admission of weakness slipped in out of caution; it is the subject.

Whoever sets the denominator holds the result

The most serious objection runs as follows: CUD is not a measure but an instrument of policy disguised as a metric, since the governance that sets minimum quality, the remediation budget and the maturation window controls the count of useful decisions, hence the result. An executive who believes he is buying the true cost per decision is buying a negotiated number. The objection lands, and proves less than it believes. CUD remains a measure; it is simply not a natural one, it is conventional and governed. EBITDA, RAROC, full costing and capital allocations form the same family, without sharing the same status, and institutions are nonetheless steered with all of them. Conventionality is not the hidden vice of CUD, it is the ordinary regime of all management accounting. What CUD adds is to make the trade-off on thresholds explicit and quantified, where cost per token buries it under an appearance of technical neutrality.

It remains that the partition of the numerator is never purely analytical. As an archetype, and not as an accounting truth, O is mostly visible in IT or platform budgets, K in the governance and risk functions, H in the business line or the central balance sheet; reality blurs that split. A cost per decision crossing several functions will be arbitrated at its seams, and the power to set the thresholds is a real power. From this follows an anti-Goodhart rule, since a steering target becomes manipulable through its definitions: any modification of the classes, the thresholds, the windows, the remediation rules or the imputation keys produces an explicitly documented break in series. A metric whose definitions are not governed ends up measuring the accounting creativity of those it evaluates.

The seat-kilometre, provided its limit is imported with it

Commercial aviation offers the best precedent for a production unit, and the limit must be imported with the analogy rather than after it. An airline does not measure its performance by the litre of kerosene burned but by the cost of the available seat-kilometre, computed under a strict constraint of safety and continuous maintenance of fitness for use. CUD is the equivalent for decision systems: a production unit priced under a quality constraint, not a raw input. The limit is that the seat-kilometre, without being perfectly homogeneous, tolerates an aggregation that the decision does not tolerate without strict segmentation into comparable risk classes. An aggregated CUD without that segmentation is a number without a referent. The consequence is a publication rule: no portfolio CUD is published without the distribution by class and without a tail indicator, failing which a critical class, low in volume but heavy in risk, bleeds without the mean signalling it. A word finally on what CUD is not: it compares architectures, arbitrates a mode of delegation, tests a business case and detects a migration of costs, and it does not measure the value of AI. Confusing it with a measure of value would reproduce, one notch higher, the error of unit that it corrects.

An ungoverned representation chain destabilises the calculation

The economic stability of CUD rests on a condition the equation does not show: the governability of the representation chain that produces the decisions, models, embeddings, guardrails, APIs and orchestration dependencies. Let a vendor unilaterally modify its upstream model, its moderation policy or the latent space of its embeddings, and three things move together. The proofs already recorded lose part of their reach, because they attest to a system that no longer entirely exists. The cost of requalification and non-regression rises. The distribution of errors deforms, and the calibration of the risk charge with it. An exogenous variable, insufficiently framed by contract, then moves two terms of the numerator at once. The lesson is not that models must be owned: an institution needs rights sufficient to identify versions, be notified of changes, reproduce decisions, reconstruct proofs and migrate uses. The key notion is contractual and technical governability, not ownership.

The token measures a resource, legitimate for billing and for monitoring, illegitimate for judging efficiency. The commoditisation of compute does not reduce the cost of the decision as it reduces the cost of inference: it recomposes its charge toward supervision, proof and risk, and sometimes toward other payers. Scarcity does not disappear, it moves. What remains to be paid when inference costs nothing is not a compute bill, it is a proof bill.

There remains the condition named without being treated, the governability of the representations that support the calculation. That is the object of the next note, Sovereign Asymmetry, which deals with what a decision costs when its representations are not governed.

Full argument, with the decomposition of the proof base K, the actuarial calibration of the risk charge H, the portfolio aggregation rules and the analysis of regime dominance, in the PDF below (11 pages, four annexes).

Read the document