Article — Position paper · ○ Open access

An effect attributed to the layout may be nothing but disguised regularization

Representation, regularization and functional complexity: a controlled experimental disentanglement on reconstructed survival data (TRANSFORM, ZUMA-7)

Jérôme Vetillard · · Twingital Institute · 21 pages · 7 min read
🇫🇷 Lire en français ↓ Download PDF

A tensor layout declares an available geometry, it does not impose an invariance

The circulating formula is seductive and, as it stands, false: choosing a layout would amount to choosing a symmetry group. It is false because an invariance never belongs to the representation alone. Write the model as the composition of a representation and a head, f = h ∘ φ. The layout fixes φ. It makes available an order, a neighbourhood, a sharing of parameters. It says nothing about what h will do with that offer. Independent columns make permutation invariance natural without imposing it: a regression assigning a distinct coefficient to each column does not permute its parameters along with its inputs. A continuous axis makes a neighbourhood available without entailing invariance to every monotone reparametrization of time, since a spline basis has its derivatives and smoothing penalties deformed by such a transformation. The defensible formulation is more modest and considerably more solid: a layout makes certain classes of symmetry and parameter sharing cheaper than others, effective invariance resulting from the triple representation, architecture, constraint.

Three levels the word symmetry conflates, and which nothing licenses stacking

The first level is the structure of the input space: order, metric, neighbourhood, hierarchy. The layout fixes it alone, and that is the whole of its authority. The second is the functional hypothesis, smoothness, locality, additivity, proportionality, that the representation-model pair makes cheap or prohibitive. Only the third is symmetry proper, invariance or equivariance under an explicitly defined group. Conflating the three yields a doctrine that explains everything and discriminates nothing, the classic symptom of a concept grown too extensive. Holding the cut has an immediate consequence for what can be measured: the experimental unit is no longer the isolated layout but the complete representational pipeline, layout, parametrization, penalization, head. This is the same attribution requirement met in any complex socio-technical pipeline, where one never measures a component but always a chain.

The protocol rests on three constraints: public sources, a single family, a design frozen before the confirmatory run

The bench uses no individual data. Time-event-censoring datasets are regenerated from the published Kaplan-Meier curves of TRANSFORM and ZUMA-7 alone, two second-line trials in large B-cell lymphoma, by the method of Guyot and colleagues. The right term is compatible, not identical: the algorithm does not recover real trajectories, it produces a distribution that reproduces the curve. Validation re-estimates the hazard ratio on the reconstructed dataset and lands within three hundredths of the published value (0.375 against 0.359 for event-free survival in TRANSFORM, 0.730 against 0.729 for overall survival in ZUMA-7). What this proves: the reconstruction preserves the scalar that governs treatment inference. What it does not prove: preservation of the complete structure, censoring times, non-proportionality, restricted mean survival. A single model family is used, a discrete-time survival model fitted by penalized log-linear regression, so as to reduce inductive biases competing with the one under study. Expressivity budgets are frozen and hashed before any draw, and the shape of the confounder is not chosen but sampled along a regularity axis, from rough to smooth, including shapes no layout can express. The discipline is that of an applicability domain: a representation is not judged on its average performance but on the region where its prediction remains supportable, and on the rule that declares that region before measuring it.

At matched effective degrees of freedom, the advantage of continuity disappears

The raw result initially favours the shared basis: representing time by a smooth basis rather than by independent columns lowers reconstruction error, and the gap widens with the number of intervals, where one column per interval overfits. The controlled ablation takes that advantage away from continuity and hands it to regularization. At matched effective complexity, measured by the trace of the smoothing operator and reached by bisection search on the penalty parameter, a discrete representation with a second-difference penalty equals a true locally supported penalized spline: ISE of 3.5·10⁻⁴ against 3.5·10⁻⁴ on smooth truth at six degrees of freedom, 5.5·10⁻⁴ against 5.6·10⁻⁴ on break truth, with identical restricted mean survival, integrated calibration and test deviance. The choice of comparator was in fact deciding the verdict: a global degree-five polynomial, mechanically penalized on a discontinuity, inflated the gap fourfold. With a local P-spline the differences fall to around one millionth and sometimes tilt towards the discrete. The lesson is twofold and only one half should be kept: statistical significance is not practical relevance, and replacing a pro-continuous doctrine with a pro-discrete mirror image would be the same error committed in the other direction.

Under simulated confounding the verdict is null, and the null is quantified rather than asserted

The second arm does not estimate a clinical causal effect but the recovery error of a conditional coefficient in a simulated world, which is a far weaker and far more tenable claim. The reported unit is the paired difference between layouts on the same seeds, resampled by a bootstrap whose unit is the seed and not the subject. None of the eight cells survives a Holm correction, the smallest p value being 0.038 against a corrected threshold of 0.006; the paired sign tests are all non-significant; the probability that the continuous beats the flat stays between 0.43 and 0.59. The cell that was marginal before correction resists neither Holm nor the sign test, which marks it as a multiplicity false positive. A null cannot be read without its margin: the minimum detectable difference at eighty per cent power is about 0.022 in log-odds, and the data are compatible with the absence of any gap greater than 0.033, that is two to five tenths of a percentage point on a baseline hazard of five to twenty per cent. It is a null within a margin, not a proven equivalence. The recoverability filter deserves naming, since it was amended: the initial per-seed filter rejected two thirds of configurations for sampling noise rather than structural non-identifiability, the oracle standard deviation being 0.10 against a threshold of 0.05. Decoupling the structural grid from estimation and raising the number of events brings that standard deviation to 0.07 and restores retention of eighteen to twenty seeds out of twenty. The grid looks only at the generator, never at which layout wins: that is exactly the line between correcting power and settling the outcome.

Three independent generators preserve the ordering, and the surviving nuance saves the principle

The charge of an endogenous generator is answered by replication, not by protestation. The same protocol runs on three structurally independent families, a Cox proportional hazards model, a Weibull accelerated failure time model, a discrete logistic model, with a smooth confounder acting on treatment and on outcome and adjustment on the true confounder as reference. The ordering of representations is preserved from one generator to the next. Two nuances honesty imposes. The spline here retains a modest advantage on a smooth confounder, whereas discrete and continuous were equal on the representation of time: matching the trace does not equalize how well a step basis approximates a smooth target. That advantage narrows by half to two thirds when discrete grain moves from twenty to ninety-six intervals, without vanishing. Grain is therefore a lever distinct from effective degrees of freedom, and the conclusion is not that one representation wins universally: it is that any comparison of representations must control both effective complexity and functional class, failing which the gap attributed to the format mixes the two.

The serious objection is that the author fabricates the world he measures

A poorly constrained comparative simulation makes it possible to claim superiority for almost any method. The objection is the right one, and the answer is to remove the choice from the author’s hands rather than to swear to his good faith. The design follows the ADEMP standard, aims at a neutral comparison in the sense of Boulesteix and colleagues, freezes and hashes budgets before any draw, prespecifies the regularity grid so that no layout expresses it by construction, and keeps the recoverability grid blind to the winning layout. None of these safeguards proves neutrality, and this should be said without reservation: local hashing does not equal a time-stamped third-party deposit, the causal generating mechanism is written in the language of the estimation family, budget matching is parametric and not functional, and forty seeds around four curves from two trials do not make forty clinical contexts. Strong validation, generators built by an independent team, an external challenge dataset, plasmodes on real covariates, replication by a second implementation, is not an optional extension but a condition of external validity. That the causal result is null, with no layout favoured, is at least consistent with a bench that does not lean. It is not proof of it.

The principle travels even where the result does not

What the bench establishes takes few words: at identical nominal information, the representational pipeline modifies the functional class and the regularization available to the estimator, and in this survival bench the apparent advantage of a continuous representation is explained by effective degrees of freedom rather than by continuity itself. Layout as an invariance hypothesis in the strong sense remains a doctrinal formula these results do not support, and it is better declared than implied. What remains is a design principle, more transportable than the local result: any comparison of representations must explicitly match or characterize effective functional complexity and induced penalization. That principle exceeds survival analysis and holds wherever representations of unequal capacity are opposed, tabular data, text, image, graphs, exactly as an applicability domain measures a proximity and not a capacity, and as a benchmark measures a performance without governing its provenance. It joins, finally, the doctrine of projection instruments: to represent is not to reproduce, and a representation is judged only by the domain in which it remains substitutable. Before attributing an effect to a format, match the complexity. Otherwise one calls geometry what was only smoothing.

Read the document