Article — Position paper · ○ Open access

To represent is not to reproduce

Projection, validation, and the domain of substitutability of synthetic cohorts

Jérôme Vetillard · · Twingital Institute · 9 pages · 5 min read
🇫🇷 Lire en français ↓ Download PDF

A tabular generator is a projection instrument, not a patient factory

The common vocabulary speaks of synthetic patients, and it misleads from the first word. A tabular generator does not imitate individuals; it learns, from a finite sample, an approximation of the joint distribution of a set of heterogeneous variables, then draws new realizations from it. The produced cohort is not a copy of the training cohort, no more than a second throw of the die is a copy of the first. The right image is not imitation but projection: a transformation that preserves certain properties of the clinical survey, discards others, and reorganizes the rest. Validating such an instrument is therefore not measuring a resemblance, but delimiting the zones where this projection remains substitutable for the observed distribution, and those where it becomes speculative. TweenMe is one of the grounds where this delimitation stops being a stylistic clause and becomes an engineering constraint.

Why synthesizing a clinical table is not synthesizing an image

The asymmetry is structural, and it forbids importing the recipes of the visual. An image possesses a geometry given by the sensor: a grid of pixels, invariance under translation, a natural neighborhood between nearby points. A clinical table has none of this. Its variables are heterogeneous, its marginals often multimodal, its categories strongly imbalanced, and no adjacency metric pre-exists the analysis. Synthesizing an image exploits a geometry; synthesizing a clinical table first requires fabricating one. The consequence deserves to be stated plainly: a tabular generator learns an implicit metric as much as a distribution, and metrics cut for images poorly diagnose its failures. The tabular is not a variant of the visual; it is a regime of its own, with failure modes of its own.

Coverage is not resolution, and the clinic pays for the confusion

The estimated support has an extent; it also has a resolution, and the second axis decides everything. The extent says where the map exists; the resolution says with what fineness it informs. This fineness is not a matter of density but of local information: a thousand nearly identical patients resolve a region poorly, whereas thirty sufficiently varied patients resolve it well. Hence a formula that holds the rest together: a region can be covered without being sufficiently resolved. Coverage reassures; resolution decides. Two failures that aggregate distances confound must then be separated. Extrapolation is a coverage failure: one exits the support. Under-resolution is a grain failure: one stays within the support without density enough to constrain it. Both produce observations plausible in appearance, and that is precisely what makes them dangerous, all the more so as the conversational fluency of modern tools conceals the transition between estimation and extrapolation.

Every generative architecture is an inductive bias, not a progress over the previous one

Families of architectures are presented as a lineage of refinements. The reading is misleading: all aim at the same target, a distribution, but each builds it with a different operator. A Bayesian network factorizes an auditable joint, a copula dissociates marginals from dependence, CTGAN fights continuous multimodality and categorical imbalance, TabDDPM learns the distance that separates data from noise, a causal generator preserves the effect of an intervention where preserving P(X) does not preserve P(Y | do(X)). None is better in the absolute; each chooses what the projection preserves, discards, and reorganizes. Inductive bias is not a defect to be corrected: it is the formalization of the assumptions that the projection imposes when observations become insufficient. Hybridization is no exception, under one condition: grafting attention or a latent diffusion has value only if each component answers an identified failure mode. On small clinical cohorts, attention grafted with no failure to correct mostly learns noise and produces a matrix that impresses committees without explaining anything.

Diffusion dominates at the center and stays undemonstrated at the edges

The serious objection must be stated at full force: if diffusion models now dominate most tabular rankings, why treat architecture as a second-order decision? The answer does not contest the benchmarks. They measure, better and better, fidelity, diversity, privacy, utility. They nonetheless aggregate these performances over average regions, and remain poorly equipped to qualify a conditional substitutability in a rare clinical margin. The dominance of diffusion is real at the center and undemonstrated at the edges. The leaderboard is true and beside the point. What is decided in the clinic is decided precisely where the ranking falls silent.

Fidelity is not substitutability, and TSTR does not grade a generator

Here is the result. An objection appears natural: if the cohort faithfully reproduces the distributions, why would it not replace the real cohort? It confuses the quality of a representation with the validity of a reasoning built upon it. Two road maps illustrate it: the first reproduces the relief perfectly but omits secondary roads, the second simplifies the relief but preserves the whole network; the first resembles the territory more, the second leads more surely to destination. Fidelity approximates the preservation of P(X); the conclusions depend on P(Y | X). Train-on-Synthetic, Test-on-Real does not deliver a grade to the generator: it qualifies a representation relative to a family of tasks, and its result depends simultaneously on the generator, the task, the downstream model, and the evaluation protocol. Two generators can invert their ranking depending on whether one trains a survival model, a classifier, or a causal estimator. Substitutability never characterizes a generator; it characterizes a representation relative to a family of decisions. The thesis is falsifiable, and it must be said so that it remains scientific: a strong and stable correlation between fidelity metrics and operational substitutability, measured across varied tasks, would refute it.

Validation is not a certificate, it is a domain to maintain

A conditional proof holds for a population, a period, a family of tasks, a protocol; modifying any one of these components modifies the validated object. A cohort validated on a breast cancer does not become valid for a bronchial cancer, and a validation on 2015 data does not necessarily survive the introduction of a new therapeutic strategy. The deep reason is twofold, and the second is rarely named: not only does the population drift, but the measurement instrument drifts with it, since nomenclatures, diagnostic criteria, and coding practices are renewed. An imperfect sample is corrected by more data; an imperfect instrument is corrected only by knowledge of its biases. The consequence is operational, and industry regularly underestimates it: the real cost of a generative platform is not the production of a cohort, it is the maintenance of its domain of validity, through monitoring, revalidation, recalibration, and versioning. This is the logic of QSAR applicability domains in toxicology: a model is not valid because it predicts well in general, but because its domain of validity is known, documented, and respected. TweenMe implements this doctrine: the product does not seek to optimize a generator, it seeks to industrialize the maintenance of the domain of validity.

Read the document