It is a capacity under constraint: workstation saturation, error correlation, and the measured sensitivity of oversight
The French doctrine of the garantie humaine and the human oversight of article 14 of Regulation (EU) 2024/1689 do not designate the same object, and public debate has been conflating them for five years. The confusion has a practical consequence: we keep requiring human oversight without ever measuring whether it is one.
The first doctrine formed around article L. 4001-3 of the French Public Health Code, created by article 17 of Law no. 2021-1017 of 2 August 2021, and around the opinions that surrounded it, from CCNE opinion no. 130 in May 2019 to the joint opinion CCNE 141 and CNPEN 4 of November 2022, which recognises human guarantee colleges. It is multidisciplinary and deferred: it organises a collective gaze upon a device, upstream and downstream of its use.
Article 14 belongs to another logic. It requires high-risk systems to be designed, human-machine interfaces included, so that natural persons can effectively oversee them during the period in which they are in use. The regulation demands neither decision-by-decision review nor real time. But in the class of systems at issue here, serial decision support at sustained cadence, the requirement becomes individual and synchronous: it must be exercisable at the workstation, within the useful time of the decision. The first doctrine was promised while the second device was funded. This text bears on the second, and on that class alone: it does not hold for rare decisions with high unitary stakes, where deliberation time exists by construction.
Paragraph 4 lists capacities that sort into three families. To act: disregard the output, refuse to use the system, interrupt its operation. To understand: know the system’s capacities and limits, interpret its output correctly, monitor its operation. And a third, of a different nature, remain aware of one’s own tendency to rely excessively on what the system produces. The first belongs to ergonomics, the second to training, the third to metacognition.
None of this is a property of the system. These are capacities exercised by a person, at a workstation, under a load. The regulation designs a device and never instruments the capacity.
What saturates a workstation is not the flow of decisions but the product of that flow by the quantity of judgement each decision demands, related to the cognitive capacity actually available. Two hundred files a day of which one per cent carry real ambiguity are controllable. Twenty files a day each requiring fifteen minutes of independent analysis may not be. The same workstation, the same person, two opposite regimes.
S = (λ × J) / C, where λ is the frequency of decisions submitted to the workstation, J the average judgement demand per decision, and C the effective judgement capacity available. Below 1, the workstation holds a cognitive reserve and oversight can be what it claims to be. Approaching 1, it becomes fragile with nothing signalling it. At or above 1, systematic independent review is not degraded, it is structurally impossible, and what subsists under that name is a ratification procedure.
This is a design heuristic, not a calibrated instrument borrowed from cognitive ergonomics: neither J nor C is directly measurable, and writing it this way does not make them so. The ratio is written on averages, while its distribution matters more: a workstation sustainable on average saturates under clusters of high-demand cases, precisely when oversight would be most necessary. What the notation brings is a displacement of the question: stop asking how many files the operator processes, start asking how many judgements the workstation demands and how many it can produce.
It also exposes a consequence the literary formulation concealed. Introducing an AI system acts on all three terms, and not in the same direction: it raises λ, since that is generally the point of the deployment; it lowers J on simple cases, which it handles well; it raises J on residual cases, precisely those it handles badly and leaves to the human; and it lowers C over time, through deskilling. A device can therefore degrade its own oversight without any of its indicators moving.
As the ratio approaches unity, oversight does not disappear, it changes in nature: its output no longer depends on the content of the case but on the configuration of the workstation. Under load, oversight converges towards the action the workstation makes free. An alert is an interruption, the free action is to dismiss it, load produces mass rejection. A review station is a validation station, the free action is to approve, load produces ratification. Same mechanism, two settings of the default. The design variable is not vigilance, it is the setting of the default action and the distance separating the operator from saturation.
Effective oversight presupposes a judgement capacity, an epistemic independence, an effective authority and a sustainability of disagreement. These do not add up: in the heuristic proposed here they behave multiplicatively, H = f(S) × I × A × D with f decreasing. A value close to zero on a single term annuls the contribution of the other three. That is a property of the model, not a measured regularity, and the form of f remains unspecified: nothing today tells us whether the degradation is gradual, threshold-based or abrupt.
Authority is not the button. An operator may have the time, detect the error, have a perfectly ergonomic override command at hand, and exercise no effective oversight if disagreement triggers a hierarchical justification, degrades their indicators, or may be held against them later. The regulation names it, in recital 48 and in paragraph 5, in a provision it applies to a single case.
Sustainability of disagreement is distinct from authority, with which it is readily confused. Authority is the permission to deviate; sustainability is what deviating consumes, and how many times a day one can consume it. In the ordinary regime, ratifying costs nothing while deviating commits, documents, slows down and is seen. The asymmetry is not a law of nature: it reverses as soon as an organisation demands accounts for ratifications as much as for overrides. That does happen, generally before a commission of inquiry.
Epistemic independence is the most poorly understood condition, because it is reduced to information. An operator may detect an error from the same data through a different representation or mental model: the independence required may be informational, methodological, temporal, organisational or causal. The real problem is not data sharing, it is error correlation. Safety engineering has known this for decades as common-mode failure, and the vocabulary of successive barriers used here is that of James Reason (Human Error, Cambridge University Press, 1990).
Two readers trained in the same school do not necessarily constitute two independent barriers. Two models sharing training data and architecture do not either. And the limiting case deserves a name, because it is becoming the implementation norm: an operator who verifies an output on the basis of the explanation supplied by the system that produced it is not a second barrier, it is a pseudo-redundancy. The explanation improves understanding and simultaneously frames the reasoning of the one receiving it; its provenance belongs to the architecture of the barrier, on the same footing as its content. The value of a second barrier does not depend on its own performance, but on the correlation of its errors with those of the first.
The legal objection must be named here, because it is the strongest one. Article 14 imposes no cognitive result: it requires the provider to make oversight possible by design, and the deployer to entrust it to persons with the necessary competence, training and authority. To reproach the regulation for not guaranteeing vigilance amounts to reproaching it for not being what it never claimed to be. The objection is correct, and that is exactly where the institutional problem lies: an obligation of designed means, uninstrumented, produces a documentary compliance whose effectiveness no one, including after the fact, can establish. The text is defensible; the use expected of it is not. The tension sharpened with Regulation (EU) 2026/1744, in force since 27 July 2026, which rewrote article 4: support the development of AI literacy among staff, without guaranteeing a determined level for anyone. The design requirement remains owed; the horizontal means has become optional in its intensity.
The aviation analogy circulates in both directions, and generally badly. It does not prove that the human is the weak link: the very category of human error partly reflects how socio-technical systems attribute causality and responsibility, which Madeleine Clare Elish described as the moral crumple zone (Engaging Science, Technology and Society, 2019, vol. 5, pp. 40-60).
What aviation does establish is more useful. Lisanne Bainbridge formulated it in 1983 in Ironies of Automation: automation leaves the human the tasks he performs least well, prolonged monitoring and recovery in abnormal situations, while degrading through disuse the very skills he will need the day automation fails. That is precisely the transient regime of the saturation ratio: the numerator rises where it is most costly, the denominator falls in silence. The sector that invests most in human vigilance has been documenting its degradation for forty years. This is an a fortiori argument, not a transposition: constrained traffic, selected operators, institutionalised feedback, nothing comparable with a workstation processing scored files on a production line. And aviation did not obtain its safety by demanding vigilance: it stopped depending on it, through flight envelope, redundancy, enforceable checklists, dual signature and recorders. In the whole regulation, a single place proceeds this way, article 14 paragraph 5, which requires two natural persons for remote biometric identification. One case, not a method.
The Dutch childcare benefits affair shows what unmeasured oversight becomes. Amnesty International describes a scoring system, nationality used as a risk factor, a manual review required after flagging, and the absence of information allowing agents to understand the score (Xenophobic Machines, 2021). The decisional architecture is that of article 14: the machine flags, the human examines. Causality there is richer than the load effect alone, and our model explains only part of the mechanism. It does illuminate the asymmetry of cost: for years, ratifying a flag cost the agents nothing; the government’s resignation in January 2021 made those ratifications retrospectively very costly, for the people least well placed to contest them. Ben Green, reviewing forty-one policies requiring human oversight, draws a conclusion more disturbing than the empirical incapacity of operators: such policies legitimise the adoption of defective systems by producing a false sense of security (Computer Law & Security Review, 2022, vol. 45). Nominal oversight does not protect the person exposed to the decision. It protects the deployment, and it exposes the operator. The mechanism is the one described in Refusing care, detecting fraud, with the difference that here the law requires it.
If nominal oversight is the problem, measurement is the answer. Still, the right thing must be measured, and the obvious measure measures nothing. Two systematic reviews established the variability of override rates for medication alerts: from 49% to 96% across seventeen publications (van der Sijs et al., JAMIA, 2006, PMID 16357358), and from 46.2% to 96.2% across twenty-three studies published between 2000 and 2019 (JMIR Medical Informatics, 2020, PMID 32706721). At the other end of the spectrum, Rosbach and co-authors measured the inverse movement in computational pathology: among twenty-eight trained pathologists, exposure to an erroneous recommendation reverses an initially correct assessment in roughly 7% of cases (arXiv:2411.00998, 2024). What the study proves: automation bias exists among experts, on a task at which they are competent, and it is quantifiable. What it does not prove: a prevalence in routine clinical practice, the sample being twenty-eight and the protocol experimental. In the same study, integrating AI significantly improved overall performance.
The two series do not add up: the first pertains to under-reliance, the second to over-reliance. Together they establish that a high override rate does not prove effective oversight, and that a low rate does not prove a reliable system. The JMIR review goes further by classifying overrides according to their appropriateness, with a range from 29.4% to 100% depending on alert type: two departments displaying the same rate may, in one case, exercise pertinent oversight and, in the other, produce noise. The rate does not separate them; the qualification of each override does.
Hence the reformulation, which must be bounded to hold. Article 14 assigns several functions to oversight, and only one lends itself to this treatment. When oversight is assigned the function of detecting and correcting erroneous outputs, it becomes itself a classifier of those errors. And a classifier is evaluated by cross-tabulation with the true state.
| AI incorrect | AI correct | |
|---|---|---|
| Human overrides | useful override | unnecessary override |
| Human ratifies | error let through | correct ratification |
From this table follow the four quantities that really describe an oversight device: sensitivity, specificity, predictive value of disagreement, predictive value of ratification. The framework is the banal one of any diagnostic test; its application to the oversight device is the central proposal of this note, and it moves human oversight from the status of a disposition to be documented to that of a function whose performance is measured. One precaution guards against naive use of the table: an output error is not an incorrect decision, and an incorrect decision is not a harm. The sensitivity of oversight measures the performance of the oversight layer, not the final reduction of risk.
The serious difficulty remains: in most high-risk systems the true-state column is unavailable at the moment of decision. Responses exist and are known to quality systems: deferred ground truth, independent double reading on a sample, expert adjudication, reference cases injected into the flow, systematic review of disagreements. None is free; all are cheaper than demonstrating after the fact that a displayed oversight was not one. This is also an evidentiary requirement: in the SCHUFA Holding judgment of 7 December 2023 (C-634/21), the Court of Justice did not ask who signed but what determined, and it recalls at paragraph 67 that the controller must be in a position to demonstrate compliance. An organisation that does not know the sensitivity of its oversight holds no measure of the proportion in which machine output is effectively corrected when it ought to be. The draft standard prEN 18229-3:2026, prepared by committee CEN/CLC/JTC 21 and under enquiry since this summer, will decide whether article 14 is operationalised by measurable requirements or by a list of dispositions to be documented.
One documented device holds, and it acts on the saturation ratio rather than on vigilance. The MASAI trial, conducted on 105,934 participants in Sweden, routes examinations scored 1 to 9 to a single reading and those scored 10 to a double reading. The publications report a 44.2% reduction in reading workload (Lång et al., Lancet Oncology, 2023), a 29% increase in detection with no significant rise in false positives (Hernström et al., Lancet Digital Health, 2025), and, on the primary endpoint, a non-inferior interval cancer rate, 1.55 versus 1.76 per thousand, ratio 0.88 with a 95% confidence interval from 0.65 to 1.18 (Gommers et al., The Lancet, 2026). What these figures prove: at a reading workload reduced by 44%, the device does not degrade the result and detects more. What they do not prove: superiority on interval cancers, the confidence interval covering unity. The protocol also holds within organised screening, a perceptual task and expert readers. The human remains fully in the interpretive loop: the innovation is not removing oversight, it is using the machine to allocate a scarce resource differentially.
Hence two levels of response that must not be conflated. At the transactional level, three patterns: measure, that is, know the sensitivity of the oversight and make it an auditable asset; reduce the judgement demand to bring J back to the level at which judgement exists; or design the system so that the human does not have to compensate for a structural weakness, which amounts to declaring the device for what it is. At the institutional level, something else: Green does not conclude in favour of better human oversight, he concludes in favour of shifting the centre of gravity towards institutional oversight, meaning admission of the system, populations concerned, thresholds, conditions of use, post-deployment monitoring, conditions of suspension. Govern, admit, operate, oversee, monitor, suspend. Transactional oversight occupies one place in that sequence and only one: a barrier of last resort at the level of the decision, not the safety architecture itself. Treating it as such is exactly the error the regulation encourages without prescribing it, and it is why you only govern what you can still redirect.
Five limits, assumed. The class of systems targeted is narrow and does not cover rare decisions with high unitary stakes. The empirical corpus on override rates comes overwhelmingly from healthcare and from rule-based alerting systems; it holds as an argument of mechanism, not as a transposable measure. The saturation ratio is a heuristic, not a calibrated instrument. The proposed measurement has a feedback effect: a device evaluated on its disagreement rate will produce disagreement, which is treated only by the whole table and not by its first row. Finally, and most seriously, ground truth is not always ontological: in recruitment, in credit, in social protection, the right decision is normative and often counterfactual. The table remains a useful evaluation framework but ceases to be a directly measurable matrix, and the threshold between the two situations is not technical, it is political.
Stated in the vocabulary of safety, this yields a requirement harder than the regulation’s. If M is an error of the system, H the failure of oversight to correct it and B the failure of downstream barriers, the probability that a failure traverses all three is written P(M ∩ H ∩ B) = P(M) × P(H | M) × P(B | MH). This writing presupposes no independence; it makes the dependencies visible and requires that they be estimated. It forbids one thing only, and it is common practice: replacing these conditional probabilities with marginal rates measured separately, then multiplying them.
The pertinent regulatory question is therefore not whether there is a human, but what risk-reduction function is attributed to that human, what their probability of failure is under real operating conditions, and what independent barriers subsist when they fail. Posing it this way avoids making human oversight the new regulatory deity after having dismantled the old one. Human presence is not a safety barrier: it becomes one only if its function is defined, if its failure modes are measurable, and if its failures are not correlated with those of the system it is supposed to control. Oversight that never produces disagreement is not necessarily oversight. Oversight whose sensitivity is unknown is not even evaluated.
Doctrinal notes and explorations on AI in regulated systems. Once or twice a month. One-click unsubscribe.