Article — Position paper · ○ Open access

Representation Preconditions Sample Efficiency

At provably equal information, the choice of head axis decides what the machine holds as near, and the resulting gap is measurable before any training

Jérôme Vetillard · · Twingital Institute · preprint submitted to TMLR · 14 pages · 11 min read
🇫🇷 Lire en français ↓ Download PDF

Two tensors. The same bytes. Two learning outcomes. A longitudinal cohort is laid out as [patient, time, variables] or as [time, patient, variables]; the two layouts are mutually reconstructible by a deterministic bijection, verified by round-trip exact equality across configurations, so neither holds one bit more than the other. At equalized parameter budget, on the same data, the same transversal task reaches 0.997 AUC under one and 0.715 under the other. This is not a gap in content. It is a gap in ease, and ease is not a property of the model alone. The condition under which this happens is narrow enough to be named, measurable before any training, and refutable by an observation one can specify in advance.

A bijection preserves information, it does not preserve ease

The word information covers three quantities that practice conflates and that this result separates. The first is Shannon information: two serializations tied by a bijection carry the same bits, each a lossless function of the other, and on that plane there is nothing to argue about. The second is algorithmic accessibility: the same bits can be arbitrarily more expensive to exploit under one layout than under the other. The third is statistical efficiency: the target may be a minimal sufficient statistic in one representation and a deeply entangled function in the other.

The result lives entirely in the gap between the first sense and the two that follow. The point of the bijection constraint is not that it is elegant, it is that it is privative: it forbids the lazy explanation that one layout simply carried more data. When information is held fixed by construction, whatever still moves can no longer be charged to content. A bijection preserves information. It does not preserve ease.

An ontology, in practice, is the choice of head axis

What data engineering calls the ontology reduces computationally to one decision: which axis is treated as the exchangeable unit, the one over which i.i.d. batches are formed, and how the remaining axes nest. De Finetti formalized this in 1937 without putting it in these terms: to declare observations exchangeable is to posit a latent model, and choosing the exchangeable unit fixes the object whose repetition is assumed.

The predictable objection deserves to be stated better than an opponent would state it. A reader from knowledge representation will say that a tensor layout is not an ontology in the sense of entities, relations and types. He is right about vocabulary and he misses the mechanism. The decision that carries the name ontology inside a data project is, at execution time, a choice of layout, and it is that choice, not a vocabulary, that is shown to act on learning. Nothing is claimed beyond it: representation is operationalized here as the choice of grouping axis, and the article refuses to derive any thesis about taxonomies from it.

The consequence is a correction to a causal chain most teams take for granted. It is not Architecture → Function. It is Architecture → Representation → Function. The same encoder, trained on two serializations of one dataset, learns two different functions, because the representation has already fixed the proximities within which the architectural bias operates. A representation preconditions, in the strict sense used in numerical optimization: a preconditioner does not change the solution of a system, it changes the ease of reaching it.

The crossover survives a fixed encoder, and that is the only reason to believe it

A raw crossover proves nothing as long as each ontology receives the encoder tailored to it, since representation and encoder then remain confounded. The control that settles the matter is a single encoder applied two ways differing only in the axis over which it pools, by slice against by patient. It decomposes the transversal gap: 0.282 splits into 0.178 of representation and 0.104 of encoder. Slightly more than a third of the headline figure was a pairing artifact; the larger part was not.

Three controls complete the bundle, and the second is the most instructive precisely because it is negative. On a representation-neutral generator the trained gap falls to +0.001 individual and +0.000 transversal, far below the pre-registered refutation threshold of 0.05: the pipeline fabricates no gap in the absence of an asymmetric mechanism. On that same null data the model-free metric nonetheless still favors time-first by +0.19 while the trained models sit at parity, which is not a leak but the metric-versus-learnability distinction caught in the act: a structural gap yields a performance gap only when the disadvantaged model cannot reconstruct what it lacks. Finally, a k-nearest-neighbor probe, training no model, reproduces the crossover at 0.998 against 0.740. The phenomenon precedes any optimizer.

Honesty requires the counterweight. The crossover is asymmetric: the transversal advantage is clean, around 0.28; the individual advantage is modest, around 0.02, and a ten-seed bootstrap of the pre-specified comparison gives +0.033 with a 95% confidence interval of [−0.004, 0.071], which straddles zero. That leg is fragile. Publishing it fragile is better than publishing it rounded.

Neighborhood purity is measured without training, which is what makes it an instrument

A representation induces a metric, hence a neighborhood graph. Neighborhood purity is the average fraction of the k nearest neighbors sharing a point’s label: a property of the pair (metric, label), computable with no trained model, on the data exactly as served. It is that absence of a model that makes it usable upstream of a budget rather than after one.

Parameterizing tasks by a knob that interpolates the label unit from individual to transversal, the model-free alignment gap and the performance gap of a genuine learner swing together from negative to positive, and the crossover becomes a zero crossing at the same point: r = 0.996, slope 0.84. A correlation that high on a smooth sweep invites suspicion, and the caveat belongs in the text: 21 points at 7 task levels, naive bootstrap [0.995, 0.998], cluster bootstrap resampling whole task levels [0.994, 0.999], permutation p < 0.001. The problem is not significance, it is interpretation, since alignment and performance are both driven by the task family. That is exactly why the causal ablation becomes necessary rather than decorative.

The bundle replicates across four worlds differing in observation modality: longitudinal survival, a continuous panel with Student-t innovations, a bipartite matrix with no time, and a stochastic block model graph learned by message passing. The fixed-encoder crossover gap there is 0.178, 0.207, 0.108 and 0.104, the alignment-to-gap correlation between 0.996 and 0.999, with permutation p < 0.001 everywhere. The generality claimed is of the sign, not of the magnitude, and two limits travel with the result: the four worlds share a two-scale latent skeleton, so their independence is of modality rather than statistical, and they run through a single implementation, so a common pipeline artifact is not excluded.

Prediction is not causation, and the invertible rescaling makes the difference

A correlational predictor does not say which way to pull the lever. To move from prediction to cause, the intervention acts on the metric itself, at fixed label and fixed information, by an invertible diagonal rescaling of the discriminative block. An invertible map removes no bit and still changes the Euclidean metric, hence purity, hence learnability for a neighborhood-sensitive learner. Purity and the performance of a kernel machine then rise together, r = 0.942, above the pre-registered threshold of 0.7, and the learner’s error turns out to be affine in impurity, R² = 0.887, as a smoothness bias predicts.

Two scope conditions are declared rather than implied. Diagonal rescalings form a subfamily of the set of possible metric deformations. And the intervention acts on constructed coordinates, which establishes causality along metric → neighborhood → performance, not end-to-end from the representation, whose first link rests on construction and on the model-free reproduction above. The two extremes of the sweep are exactly the two ontologies: the choice of ontology is therefore not a binary alternative, it is a point on a continuum of metrics.

A ceiling is not a slowdown

Sweeping cohort size under the fixed encoder delivers the result hardest to swallow for anyone hoping a bad layout is recoverable with more data. Time-first reaches AUC above 0.90 by N ≤ 200 and stays near ceiling. Patient-first plateaus below 0.85 and does not close the gap over two orders of magnitude of sample size.

The vocabulary must stay exact, because two distinct claims hide under one word. The sample-efficiency claim says the better-aligned representation reaches a common threshold with fewer samples; it holds where both serializations converge. The representational-ceiling claim says the misaligned representation does not reach the other’s performance at all; it holds where debiasing is impossible. They share a cause and they are not the same statement. The criterion deciding which applies is checkable: can the misaligned grouping, in principle, recover the label. On a bipartite task whose label is a genuine group aggregate but carries no inaccessible nuisance, the two representations converge to within 0.001. The ceiling is therefore not a property of misalignment in general; it is the consequence of a misalignment that renders a nuisance unreachable.

Minimality is refuted, and publishing that beats rescuing it

Neighborhood purity alone explains R² = 0.992 of the performance gap and adding separability improves it by only 0.006: it is a sufficient statistic. It is not a unique one. Separability alone reaches 0.997, marginally better, and the partial correlations lean against purity. The data do not crown purity, and the text says so against its author’s initial expectation.

Broadening the competition sharpens the diagnosis instead of diluting it. Four label-aware structural measures, purity, separability, mutual information and a geometric margin, all predict the gap between 0.99 and 1.00. Two label-agnostic geometric descriptors, effective rank and local intrinsic dimension, predict nothing at all, r = 0.00, because they do not even distinguish the two serializations, which are re-centerings of the same points. The active quantity is label-aware class structure in the induced metric; purity is a sufficient and convenient estimator of it, not the fundamental variable. A geometry blind to the question sees nothing, however refined the descriptor.

The predictor names its failure regime, which is the condition for using it

An indicator claiming to hold everywhere holds nowhere in governance, because no stopping criterion can be written with it. Two pre-registered tests bound this one. On a task swept from linearly separable to XOR-structured at fixed representation, purity stays moderate around 0.81 at the XOR pole and a radial-basis-function kernel machine tracks it faithfully, r = 0.99 across the sweep, while a global linear learner collapses to chance, 0.50, despite that purity. The predictor therefore characterizes learning dominated by the induced neighborhood: nearest neighbors, kernel machines, empirically tree ensembles. A globally-biased learner on a locally pure but globally nonlinear label escapes it, and a formal characterization of that validity domain remains to be written.

The pre-registration assumed the failure would strike the kernel learner. The opposite occurred. The inversion is reported as it stands, which is the only reason to write a pre-registration at all. The second test injects information loss into one serialization, so the two are only approximately reconstructible: the alignment-to-gap correlation degrades gracefully and stays above 0.99 up to substantial corruption, because the indicator is empirical and measures the purity of the representation actually supplied. Exact bijection is a clean sufficient condition, not a fragile requirement, which moves the result out of the demonstration regime and into engineering.

Learnability is bounded by the representation the model is plugged into

The Institute established elsewhere, on a toxicological control device, that the information available to a governance device is bounded by the representation it is plugged into. The present work supplies the symmetric half of that statement, and supplies it measured: the learnability available to a model is bounded, likewise, by the representation it is plugged into. An audit cannot recover a property the featurization destroyed; a model cannot recover, with more data, a proximity the serialization undid. Both statements descend from the same elementary ordering: a quantity computed downstream of a channel knows no more than the channel did.

The consequence is positional, as it was for the validity port. The repair is not to enrich the model but to measure earlier. Neighborhood purity is computed on the data as served, before any budget is allocated, and it gives the sign of the gap before a GPU is reserved. A team that measures it does not gain a better model; it gains the right not to fund the wrong layout for two quarters. This is a change of plug, not a gain in richness.

The engineering heuristic fits in four ordered moves: identify the unit of decision, build the ontology whose elementary objects coincide with it, verify that it preserves the relevant dependencies, adapt the architecture only then. On a digital twin generation platform such as TweenMe, where one cohort feeds both individual and population-level questions, this ordering stops being a stylistic clause: it decides what is modelable at constant budget, and it decides it before the first parameter is initialized. The instance establishes nothing on its own; it indicates where the constraint bites. An ontology is not an unconditional gain, it is a bet on alignment, and that bet can be lost. The most instructive case remains the one where the richer representation learns the least, because its richness is orthogonal to the question asked.

Domain of validity

The thesis holds for learning dominated by the induced neighborhood, which covers a broad but not total share of deployed machine learning, and it does not hold for learners whose bias is global or symbolic. The demonstration is synthetic, completed by two real longitudinal panels, dietox and BtheB, where the predictor holds at r = 0.983 and r = 0.996 with a zero crossing in both cases. The real panels confirm the downstream, correlational half of the chain; they do not confirm the causal half, since the representation-to-metric intervention remains entirely synthetic. Saying otherwise would be extrapolation.

Four further limits belong to the result rather than to its margins. The four worlds share a two-scale latent skeleton, and a generator breaking that structure remains to be tested. The ablation covers only a subfamily of metric deformations and stays local to the metric. Minimality is refuted, not established, and the formal derivation is validated in form rather than written as a bound. Finally, because the two poolings induce different input and minibatch distributions even at a fixed encoder, a residual distributional confound cannot be fully excluded, and the task family is a constructed interpolation with no guarantee that genuinely independent tasks would reproduce the same continuity.

The observation that would refute the thesis is nameable, and that is what makes it a thesis. A pair of serializations in bijection, on a neighborhood-dominated task, exhibiting a clean alignment gap and no performance gap, would reduce it to a coincidence of the task family. It was not observed here, including on the null generator, where the performance gap correctly vanished.

At constant and provably equal information, the representation fixes the proximities a learner uses, and therefore what it can learn from the data it has. The question is not whether your layout contains the information. The question is whether your layout makes it near.

Full argument, the four-world protocol, the fixed-encoder decomposition, the causal ablation, the broadened mediator competition, the pre-registered failure regime, the real-panel validation and the generator equations in the PDF below (14 pages, EN version of TMLR submission v3.3).

See also: To Represent Is Not to Reproduce · An Applicability Domain Measures a Proximity, Not a Capacity · The Model Was Never the Object · Learning What Cannot Vary · Governing Trajectories, Governing Invariants · Clinically-Informed Neural Networks

Read the document