What a comparison of representations establishes, and what it cannot yet establish: preregistered evidence at matched effective degrees of freedom
The previous instalment of this series closed on a design principle: before attributing an effect to a format, match the complexity, otherwise one calls geometry what was only smoothing. This one runs the experiment that the principle demands, and the experiment declines to return the comfortable answer. At matched effective degrees of freedom, the differences between representational transformations do not vanish. They shrink, they reorganize by regime, and they persist.
The result compresses into a double non-implication. Equality of effective degrees of freedom implies neither equality of the accessible function classes nor equality of the biases. Two pipelines granted exactly the same budget of flexibility can operate over different function classes and retain different performances in estimating the same causal estimand. The methodological consequence is narrow and worth stating without inflation: matching global complexity eliminates one competing explanation, a difference in global capacity budget, and eliminates only that one. It does not identify what remains.
Two questions are routinely conflated in the literature on representation comparison, and their conflation explains a large share of the disagreements. Which pipeline works best calls for comparison at optimal configuration, where imposing equal complexity would be an artefact. What share of the difference comes from the representation requires matching. This work answers the second question. Its conclusions do not transfer to the first, and anyone reading a leaderboard as an answer to both is reading two experiments in one table.
The naive question opposes time carried as an axis to time split into columns. It is ill-posed, and the earlier formulation of this research programme, at identical input information, was false in the statistical sense. The four transformations under test do not preserve the same information. The ordered grid performs a many-to-one compression. The wide format destroys part of the continuous metric. The continuous axis preserves it and adds a quadratic basis. The hierarchical form injects a structure absent from the source covariates.
Four operations are therefore conflated under a single word: arrangement, encoding, feature engineering, restriction of the function class. This testbed does not disentangle them, and saying so costs nothing but a headline. What is established is not that representation matters, a claim so accommodating it discriminates nothing, but that the construction of the feature space remains decisive after equalization of a global complexity budget. Less spectacular, considerably harder to refute.
The word symmetry covers three distinct things, and stacking them is the principal conceptual weakness of the field. The first level is the structure of the input space, order, metric, neighbourhood, hierarchy, which the transformation fixes. The second is the functional hypothesis, smoothness, locality, additivity, that the transformation-model pair makes cheap or otherwise. Only the third is symmetry proper, invariance under an explicitly defined group. Writing the model as f = h ∘ φ, effective invariance depends on both terms jointly; the transformation fixes φ and determines nothing alone. These experiments test the second level. The right name for what they measure is structural adequacy, and symmetry stays reserved for the theoretical frame.
Individual data are reconstructed from published Kaplan-Meier curves by the method of Guyot and colleagues, accepted by NICE and the HAS, and validated against published hazard ratios: 0.375 against 0.359 reconstructed for event-free survival in TRANSFORM, 0.730 against 0.729 for overall survival in ZUMA-7. The inversion operates on an aggregate curve and does not restore the covariates. What follows is a semi-synthetic simulation calibrated on published covariate margins, and the point deserves stating flatly because clinical cohort names in a methodological article spontaneously manufacture an impression of empirical validation. What comes from the cohorts is the marginal distributions and nothing else. What does not come from them is the correlations, the causal mechanism, the shape of the temporal heterogeneity, the interactions and the censoring regime. No conclusion about lymphoma, about CAR-T effectiveness or about clinical decision-making follows from this work.
A single estimation family is used, a discrete-time survival model fitted by penalized logistic regression in person-period format, which reduces the inductive biases competing with the one under study. The family is not opinion-free: it imposes an additive form, a link function, a quadratic penalty and an interval approximation. The study therefore bears on one slice of the space of questions, R × DGP at fixed M. Whether the observed differences are intrinsic to the function spaces or arise from their interaction with this particular penalty remains open.
The estimand is the marginal difference in restricted mean survival time at a fixed horizon, obtained by g-formula. The choice is not decorative. A conditional coefficient depends on the covariates included, so two transformations compared on their own coefficients do not aim at the same target and the comparison is void before it starts. Complexity is matched on the trace of the smoothing operator, the standard measure of effective degrees of freedom for a penalized linear estimator, with the penalty of each transformation solved by bisection to reach the oracle target before any reading of the bias. Code, protocol and analysis plan were deposited and timestamped before execution, and the confirmatory run verifies the code fingerprint at launch and refuses to execute if it differs. This is the difference between a preregistration and a promise: the second is a claim about the author, the first is a check a stranger can run.
Sixty conditions, two thousand replications, two sets of margins, four conditions excluded from the primary analysis and all of them under hidden confounding. On the absolute bias on ΔRMST, the continuous axis reaches 0.2147 against 0.2318 for the wide format under multiple confounding, 0.9265 against 1.0171 under temporal confounding, 0.0223 against 0.0241 under non-linear confounding, for a Monte-Carlo standard error near 0.004. The hierarchical form lands between the two under the first two regimes and collapses to 0.0850 under the third.
The decision runs on an equivalence margin of 0.033, and the margin’s origin is stated without embellishment: it is the minimum detectable difference observed during the exploratory campaign, neither a clinically grounded quantity nor a declared fraction of the RMST. Three comparisons out of seven exceed it, all under the temporal regime or against the hierarchical form under the non-linear one. The multiplicity correction rejects equality in twelve conditions out of twelve for the continuous axis under multiple confounding, where the corresponding difference is half the margin; the conclusion is taken on the margin, not on the rejection, because a difference that is real and irrelevant is still irrelevant. A margin whose provenance is that opportunistic invites the obvious objection, so the objection was run: it would take δ ≈ 0.140, that is 4.2 times the retained value and 26.8 Monte-Carlo standard errors, to reverse the verdict.
One transformation fails structurally, and the failure is a result rather than an accident. The wide format, the continuous axis and the hierarchical form reach the target trace in all sixty conditions. The ordered grid reaches it in none. Having reduced the covariates to a single score, its function class is capped below the oracle whatever the penalty. The honest description of this experiment is therefore three matched transformations and one structurally under-parameterized, whose incapacity was prespecified. Reporting it as four matched competitors would have been a small lie with a comfortable consequence.
The Monte-Carlo precision is genuine and it answers the wrong question. It bears on the error conditional on each world, and says nothing about average performance over a distribution of worlds. Four hundred worlds drawn from a declared distribution, twenty-five replications each, seven quantities that the confirmatory design treated as known becoming random: 88.6% of the variance in bias lies between worlds, and the standard deviation of performance between worlds is about 0.54, roughly one hundred and thirty times the Monte-Carlo standard error of the confirmatory run. This is the sentence the field should carry off: two thousand replications do not correct sixty worlds. It is the same failure of attribution met wherever a chain is measured and a component is credited.
Under that reading the appropriate estimand is no longer a mean bias but a probability of superiority. The continuous axis beats the wide format in 73.0% of worlds with a standard error of 0.022, beats the hierarchical form in 80.0%, and is the best of the three matched transformations in 63.0%. The average ranking is unchanged. The interpretation is not: in more than a quarter of worlds another transformation does better, so dominance here is probabilistic and not universal, exactly as a benchmark measures a relative performance without governing its provenance.
Three stress tests refuse to dissolve the advantage. Over multiples of 0.5 to 1.5 times the oracle trace, the probability of superiority varies from 0.770 to 0.790 against a standard error of 0.024, so the ranking does not depend on the level of complexity over a factor of three. A target selected by cross-validation on out-of-sample deviance, without privileged information, gives 0.760 ± 0.025 against 0.770 ± 0.024 for the oracle, which lifts the oracle reservation: matching does not require privileged information in this testbed. Across six copula families declared before execution, the gap between the non-Gaussian families and the Gaussian control is +0.021 ± 0.050, null within the noise, and the Gaussian control sits below three of the five irregular families, which makes it less plausible that the regularity of the testbed was chosen to favour the winner. Clayton is the only family below 0.60, at 0.589, and lower-tail dependence is precisely what the continuous axis expresses least well. That is not a footnote, it is an address: it says where to go looking for the counter-example.
The premise holds. Among the three genuinely matched transformations, traces agree to within 0.53% while the concentration index varies by 82%: the wide format spends its budget over 15.5 effective directions with a Gini of 0.181, the continuous axis over 10.9 directions with a Gini of 0.457, the hierarchical form over 15.8 directions with a Gini of 0.371. The standard measure of complexity says how much a model can adjust and never in which direction.
The conclusion that seemed to follow does not follow. The correlation between spectral concentration and bias is null in all three regimes, at −0.07, −0.06 and +0.05. A scalar statistic of concentration does not replace the trace. A recovery was available and was attempted: perhaps what matters is not the number of directions but their alignment with the signal. Projecting the true confounding effect onto the eigenbasis of the operator separates expressivity, the share of the true effect falling inside the design space, from retention, the share of the expressible surviving the shrinkage. Expressivity carries the within-regime correlations with bias at −0.670 under multiple confounding and −0.556 under temporal confounding, then reverses to +0.515 under non-linear confounding, which is also the regime dominated at 93% by variance rather than bias. Two regimes out of three, and a sign flip in the third.
This is reported as a hypothesis tested and not retained, which is the only useful way to report it. A better-grounded descriptor exists in the kernel-regression literature, the cumulative power distribution of Canatar and colleagues, and it was not used here; a genuine test of the conjecture should go through it and preregister it. Declaring the failed mechanism is worth more than promoting the surviving correlation, and the discipline is the same one that governs an applicability domain, which measures a proximity and not a capacity.
Twice in this work an average over heterogeneous families produced a conclusion the detail contradicts. The global correlation between alignment and bias is positive while two regimes out of three give a negative correlation. The probability of superiority is flat across sample sizes, 0.783 at n = 200 and 0.800 at n = 5,000 with a slope against log n of +0.002, while the three regimes move in opposing directions underneath: the advantage grows from 0.700 to 0.933 under multiple confounding, plateaus under temporal confounding, and decays towards chance from 0.643 to 0.500 under non-linear confounding. A flat aggregate here is not stability, it is cancellation.
The generalization is not an anecdote about this testbed. Performance aggregated over heterogeneous families of mechanisms can mask precisely the interaction the study seeks to characterize. Formally, the expectation of the loss over the distribution of mechanisms does not suffice to characterize that loss when the interaction between representation and mechanism is the object of interest. A global benchmark may then rank methods correctly and explain incorrectly why they prevail, which is the more expensive of the two errors because it is the one that transfers. This concerns any practice of comparison over heterogeneous collections of datasets, and it is the statistical face of a result the corpus has already met clinically: the same benchmark performance can produce opposite care.
The four conditions with hidden confounding are excluded from the primary analysis because the oracle itself fails to recover the truth there, with mean biases from −0.72 to −2.14. In those conditions no transformation does usefully better than another. All of them fail, and the uniformity of the failure is the point.
A transformation can improve functional approximation. It does not manufacture causal identifiability. This negative control governs the reading of the whole article, because it blocks the extrapolation that would otherwise be made towards representations, synthetic data and digital twins, where the temptation is to believe that a sufficiently good representation recovers the right mechanism. It does not, and the operational reservation is worse than the statistical one: in the real world the analyst does not know on which side of that boundary they stand. The same discipline applies to synthetic cohorts, where to represent is not to reproduce and only the domain of substitutability is defensible.
One inferential failure is declared rather than buried. The variance is correctly estimated, with a ratio of standard error to standard deviation between 1.01 and 1.03, but coverage follows the ratio of bias to dispersion: 94.8% under the non-linear regime, 70.8% under multiple confounding, 7.6% under temporal confounding. The tension must be named, since the temporal regime is the one carrying the conclusion of the confirmatory section and the one where coverage degrades most. The bias being differential between transformations, the comparison of biases remains interpretable; but in that regime none of the four would supply a usable interval in practice.
Nothing here shows that certain representations are intrinsically better. Within a preregistered framework, in a penalized family of discrete-time survival models, distinct representational transformations produce different biases despite matched global complexity, and the margin exceedances occur in a structured way across regimes. Interpreting those differences through the expressivity of the function class is a supported reading rather than an isolated one, and the spectral mechanism first proposed for them has been tested and dropped. Over a distribution of worlds the advantage is probabilistic, invariant over a factor of three in complexity, robust to the loss of privileged information and to the regularity of the dependence structure.
One measurement is declared as a limitation rather than a result, because the alternative would have been to let a number stand for something it does not measure. The protocol provided for quantifying a generative uncertainty by resampling the generator; the implementation resampled the replications already produced and reproduced the Monte-Carlo variance, ratio 1.01. It is reported as a broken instrument, not as a finding.
The priority independent test is not more simulation. The marginal return of a testbed whose authors define the structural support is decreasing, and ten thousand additional worlds would change nothing about the one dimension that matters: the space of admissible mechanisms was not randomized, only the parameters, the worlds, the covariance structures, the sample sizes and the complexity budgets. The experiment that would move the question is of another nature. A third party specifies the mechanisms, the parameters, the distributions, the interactions and the temporal dynamics, without knowing the results obtained by the representations, and the pipeline is then executed blind. No single experiment is decisive in the strict sense, but that one would bear on the only dimension these analyses cannot reach, alongside the extension to other estimation families that the R × M × DGP question demands and of which this work studies one slice.
What survives the local result is a requirement, and it is the same one that governs every other instrument of attribution in this corpus, from representation as a precondition of sample efficiency onward. Any comparison meant to attribute performance to a representation must match or characterize effective complexity, and must then say what remains unexplained after the matching. Matching is a necessary condition of attribution. It has never been an explanation.
Doctrinal notes and explorations on AI in regulated systems. Once or twice a month. One-click unsubscribe.