The system transforms the one who oversees it: on minimum independent human capacity, and on the absence of any instrument to measure it
Regulation (EU) 2024/1689 requires high-risk artificial intelligence systems to be designed so as to be effectively overseen by natural persons. It verifies that the oversight capacity is present at the moment one looks at it. It never models the evolution of that capacity under the effect of using the system. Human oversight is not an external capacity applied to the system: it is a state variable coupled to it.
The claim is not about the letter of the regulation, which obviously contains no formal model, but about its implicit ontology. The law treats supervision functionally as a constant parameter. The consequence is instrumental before it is normative: horizontal AI governance today defines neither an explicit metric for this variable, nor a standardised protocol for measuring its depreciation and its restoration.
The reproach must be bounded immediately, because its broad version would be false. Some sectors already possess partial mechanisms that serve in its stead: proficiency testing, periodic competence assessment, double reading, inter-observer agreement, clinical quality control, re-certification. What is missing is not all measurement, it is the measurement of the right thing in the right place. These arrangements were built to monitor variability between practitioners, not the drift of a practitioner exposed to algorithmic assistance. And the horizontal layer, the one that applies to all high-risk systems whatever the domain, contains none of them.
This gap is of a different nature from the existential discourse that occupies the media space. A warning about an out-of-control superintelligence indexes each of its assertions to a future, the end of next year, the end of the decade; none bears on a state measured today. That is not a failure of rigour, it is the nature of the object: a risk whose probability remains undetermined justifies precautionary work but offers no purchase for a quantitative comparison of priorities. Cognitive substitution, by contrast, is already observable, measurable in principle, and directly attached to the regulatory apparatus now in force.
The naive reading is that AI makes people worse. That is true, banal and unusable. To make it an object of governance, one must stop treating oversight competence as a single block.
| Capacity | What it enables | Regime where it is required |
|---|---|---|
| Production | Performing the task oneself | Degraded |
| Verification | Establishing whether an output is correct against a reference | Nominal |
| Recovery | Taking back control when the system fails, under time constraint | Degraded |
| Escalation | Recognising that one does not know and knowing to whom to hand over | Nominal and degraded |
Formally, H_nominal = f(C_V, C_E, A, T) and H_degraded = f(C_R, C_P, C_E, A, T), where A is formal authority and T the time available to intervene. Production capacity does not condition routine oversight: one can contradict without knowing how to redo, through detection of internal inconsistencies, calibration of uncertainty, identification of out-of-distribution cases, recourse to a second independent system, or return to primary sources. Knowing how to redo and knowing how to contradict are not the same competence. But production becomes critical again in degraded regime, when what is required is no longer to judge an output but to produce an acceptable substitute oneself.
This decomposition lifts a tension the usual formulation leaves open: one can at once exclude production from routine oversight and want to measure it as an indicator of resilience. It also extends an earlier position. In Human oversight is not a presence, supervision was described not as a state but as a throughput, saturated by the product of decision flow and judgement demand. Competence is only one factor in that throughput, and T is the one organisations compress first.
The state of the evidence reads at four levels, and conflating them is precisely what allows each camp to cite whichever study suits it.
The signal first. In August 2025, The Lancet Gastroenterology and Hepatology published an incidental finding arising from the ACCEPT trial (Budzyń K. et al., 2025;10(10):896-903). Four Polish centres had deployed a polyp detection aid at the end of 2021; across 1,443 non-assisted colonoscopies allocated by date, the adenoma detection rate fell from 28.4% to 22.4%, an absolute difference of six points (95% CI −10.5 to −1.6; p = 0.0089), with an odds ratio of 0.69 (0.53-0.89) in multivariate analysis. Nineteen endoscopists, all experienced. A correction has since added the colonoscopy indication to the covariates, without change of interpretation. What makes the study notable is not its figure but its object: the literature on assistance measures the performance of the human-machine pair; this one measures what becomes of human performance when the assistance is withdrawn. The shift of estimand is the real contribution.
The attempted refutation next. A prospective multicentre registry-based pragmatic trial (Pedersen et al., Endoscopy 2026;58(9):1003-1014) measured the same phenomenon with a superior protocol: thirteen endoscopists, 5,013 colonoscopies, three phases, before exposure, under assistance, after withdrawal. Among the non-experienced, the primary endpoint rises from 31.9% to 39.5% under assistance (adjusted OR 1.43; 1.11-1.84), then settles at 36.3% after withdrawal, with no significant difference from the initial phase (OR 1.03; 0.79-1.34). Among the six experienced operators, no significant effect in any phase.
This point is decisive and must be stated without comfort. The contradiction persists when the comparison is restricted to experienced operators alone. Nineteen Polish experts degrade, six Scandinavian experts do not budge. The candidate explanations are numerous and none is settled: the endpoint differs, adenomas versus polyps of at least five millimetres; the design differs, retrospective before-and-after versus prospective in three phases with explicit withdrawal; the power of the experienced subgroup is weak; exposure structure, case-mix, adjustments and the very definition of deskilling differ from one protocol to the other.
The state of the field confirms the heterogeneity. A systematic review published in July 2026 (Al-Anezi F. M., J Healthc Leadersh 2026;18:590498) queried four databases, identified 11,269 references and retained 29 studies under PRISMA 2020: automation bias in ten, the deskilling concern in nine, impairment of diagnostic reasoning in nine, in a literature that is predominantly observational or conceptual and too heterogeneous for quantitative synthesis. The concern is massively shared; the evidentiary apparatus is not.
There remains the level that is missing. Establishing that no longitudinal measurement of the rates of depreciation and restoration exists would presuppose a dedicated and reproducible negative search, with databases queried, search strings, cut-off date and screening log. It has not been conducted. This text therefore asserts only that such measurements have not been identified, which is not the same thing. Turning an absence found into an absence demonstrated is an academic speciality whose effectiveness stops at the first serious reviewer.
What the four levels establish together is more solid than any one taken in isolation. The interesting problem is no longer whether deskilling exists, it is that we have neither common estimand, nor common protocol, nor common horizon allowing us to decide when it exists, in whom, at what speed and with what reversibility. A discipline endowed with quality indicators among the best standardised in all of medicine cannot manage to decide whether its assistance tool depreciates its practitioners.
The mechanism, for its part, has been theorised for forty-three years. Lisanne Bainbridge, in Ironies of Automation (Automatica, 19(6), 1983), established that automation removes routine operation from the operator, the very thing that maintains the competence required for the exceptional cases automation leaves to them. With generative AI, the paradox extends to linguistic, analytical and normative tasks whose ground truth is often incomplete, costly or absent. A change of terrain rather than of nature, and the new terrain is the one where measurement is hardest.
Classical automation offloaded the gesture while leaving the measuring instrument intact: the autopilot drifts, the altimeter says so. There is not, on one side, the tools equipped with an altimeter and, on the other, generative AI deprived of one. There is a gradient, and it follows the task, not the technology.
| Task | External reference | Exposure to substitution |
|---|---|---|
| Numerical computation | strong | low |
| Testable code | strong | low to moderate |
| Document extraction | fairly strong | moderate |
| Clinical diagnosis | partial | moderate to high |
| Strategic analysis | weak | high |
| Normative reasoning, arbitration | weak | high |
This table describes an exposure, not a risk, and its ordinality is illustrative, not measured. The nuance is decisive: the available reference is not the consulted reference. A highly verifiable task produces massive deskilling if no one ever consults the reference, and the developer who accepts a suggestion without running the test is the daily example of it. Conversely, a weakly verifiable task can remain healthy if the human works against the system rather than with it.
The risk therefore depends on four terms: the rate of delegation, the external verifiability of the task, the effective engagement in verification, and the quality of the feedback received. The rate of delegation itself covers opposite regimes. Substitution, when the system does the work instead. Augmentation, when it extends the space of reasoning. Tutoring, when it produces feedback that increases competence, which does happen and must be said. Complacency, when engagement decreases without anything being learned. Two organisations using the same model in the same proportion can produce opposite trajectories. The determining variable is not the volume of use but the architecture of the interaction, which brings to light an object no one regulates: the design of delegation. It is the same displacement established in The Architecture of Feedback for reliability: at equal error distributions, two arrangements drift differently depending on how the feedback circulates.
Subject to that reservation, the dynamics can be written. Let C be the independent capacity of an operator bounded by a ceiling C_max, δ the share of cases handled without autonomous engagement, λ the rate of depreciation through non-practice and μ the rate of acquisition through practice:
dC/dt = -λ δ C + μ (1 - δ) (C_max - C)
The model remains a toy with no predictive value. It makes visible what regulatory reasoning assumes, namely that λ equals zero, and above all it brings out the right question. Not whether competence declines, but at what speed it is restored relative to the speed at which it depreciates. If τ_recovery ≪ τ_decay, the problem is one of preparation before intervention, serious but manageable. If τ_recovery ≫ τ_decay, there is hysteresis, and only then is one entitled to speak of debt. No longitudinal measurement of these two constants has been identified, neither in the systematic review nor in the protocols of the two trials that oppose each other. That is the central empirical hole in this dossier, more important than the Polish figure.
There remains the fact that δ is not exogenous, and this is what completes the transformation of the arrangement into a dynamic system rather than a decay function. A less competent operator delegates more, which makes them less competent. An expert also delegates more, but for the inverse reason, because they are better at recognising what can be delegated. The same usage statistic covers both a positive feedback loop and a rational allocation, and nothing in a dashboard distinguishes them.
To which is added a second loop, rarely formulated. In certain deployments, and presumably increasingly, human validations, overrides and corrections then serve to train, route, tune or evaluate the system. Elsewhere, this feedback is simply logged without feeding anything. It is therefore a class of architectures, not a universal property. Where the loop closes, the system progressively modifies the judge who then contributes to producing the data by which both the quality of the system and that of its supervision are evaluated. If the controller becomes complacent, the signal degrades with nothing to signal it. None of the observable quantities is then identifiable in isolation: an override rate is jointly produced by model quality, human capacity, trust, interface and workload. This is no longer deskilling, it is a co-evolution of the controller and the controlled, in which measurement itself becomes endogenous.
The relevant unit of analysis is not the operator but the organisation, and collective competence is not the sum of individual competences: C_org ≠ Σ C_i. An organisation can retain perfectly competent experts and lose its oversight capacity because they are too few, no longer in the loop, called upon too late, or present somewhere other than at the point of decision.
Above all, it depends on two stocks and not one: the experts in place, and the mechanism that produces the next ones. That mechanism runs through the tasks the organisation has just delegated. The senior remains competent for another five years. The junior never had to learn what the model now does in their place, and will therefore not become the senior they would have been. This is exactly the mechanism described in The agent does the junior’s work, not the apprenticeship, referred here to the compliance threshold: an organisation can satisfy its competence threshold today while having already lost the capacity to satisfy it tomorrow.
The signature of this phenomenon is that of invisible technical debt. Nominal performance remains constant, sometimes improves, and nothing appears in the indicators as long as the old component works. What decreases silently is resilience. In reliability vocabulary: human competence becomes a non-redundant capability whose mean time to recovery increases while nominal operations remain excellent, and no one monitors the mean time to recovery of a skill. When the seniors leave, organisational competence does not decline steadily, it goes over a cliff. A linear depreciation at the individual level produces a discontinuity at the collective level, and it is the discontinuity that makes the accident.
One last mechanism completes the picture and it is the most counter-intuitive. The better the system performs, the fewer reasons the human has to contradict it. The less they contradict it, the less they exercise their independent judgement. The less that judgement is exercised, the less available it is for the rare case on which the system fails. Yet it is exactly that rare case which motivates the supervision requirement. As AI improves, human oversight becomes less often necessary and more difficult precisely at the moment when it becomes necessary.
Ordinary vocabulary conflates two objects. An asset depreciates; a debt accumulates. Competence is the stock, cognitive substitution is its depreciation mechanism, and the debt is neither one nor the other: it is a gap between required capacity and available capacity.
D_t = max(0, C_required,t - C_available,t)
Written this way, it recalls that it possesses two sources and not one. It grows when available capacity falls, which the preceding dynamics describe. It also grows when required capacity increases, at rigorously unchanged competence, for example when a system is extended to more critical cases or to a wider perimeter. An organisation can take on debt without depreciating. The vocabulary then stabilises: the non-assisted performance test becomes an impairment test, and recurrent practice a maintenance expense. It is the direct continuation of the thesis of From certified product to continuous authority: what is certified at one instant does not remain true by inertia, and supervisory competence is no exception.
Article 14 enumerates what supervisors must be able to do: understand the capacities and limits of the system, correctly interpret its output, decide not to use it, disregard it, reverse it, interrupt it. Its point (b) enjoins remaining aware of the tendency to rely automatically on the output, designated in so many words as automation bias, the only place in the regulation where a precise cognitive phenomenon is named. Article 26(2) requires the deployer to entrust supervision to persons possessing the necessary competence, training and authority. The arrangement is coherent and rests on an unstated assumption. Remaining aware of a bias is not retaining the capacity to correct it.
The gap must be formulated with precision. To maintain that the regulation does not require the maintenance of competence would be imprudent: a contradictor would functionally reconstruct such an obligation from the deployer’s duties, from the management system or from post-market surveillance, and they would have good arguments. What is missing is not a norm, it is an instrument: the regulation requires competent supervision without defining any metric of recency for that competence, any minimum threshold of independent performance, or any periodic test establishing that the supervisory capacity remains recoverable.
The ISO/IEC 42001 standard does not fill the gap. Its clause 7.2 reproduces the generic architecture of management systems: determine the required competences, train, evaluate the effectiveness of the training, retain the evidence. What the auditor examines there are qualification files and training records; nothing calls for a measurement of the operator’s performance deprived of the tool. Conformity attests to a title; it does not attest to a present fitness.
Two sectors show the object is regulable. The European Society of Gastrointestinal Endoscopy published in 2026 a position arising from a formal Delphi consensus, defining a curriculum for the acquisition and maintenance of the competence necessary for the use of AI in endoscopy (Mori Y. et al., Endoscopy 2026;58:202-210): secure competence in standard endoscopy first, acquire the fundamentals of AI, recognise and mitigate the cognitive biases of human-machine interaction, avoid overconfidence, continuously monitor quality indicators. The sector produced in eighteen months what the horizontal layer does not contain. This is a counter-example to the broad version of the criticism and a confirmation of its narrow version.
Aeronautics encountered the same problem forty years earlier and solved it in the reverse order: it first held the licence to be a fitness, then observed that it was not. Section 61.57 of Title 14 of the Code of Federal Regulations prohibits a pilot from carrying passengers unless they have performed three take-offs and three landings within the preceding ninety days, as sole manipulator of the controls. This is not a requirement of a diploma, it is a requirement of recent and unassisted practice, enforceable, dated, verifiable. In January 2013, the Federal Aviation Administration published SAFO 13002, which states that continuous use of autoflight systems does not reinforce a pilot’s manual flying knowledge and skills and can degrade their ability to recover the aircraft rapidly from an undesired state. In May 2017, apparently judging that the message had not got through, the same administration republished the text under the title SAFO 17007, adding the word proficiency.
Aviation possesses a recency requirement, an explicit doctrine on degradation through automation, and a periodic arrangement that tests non-assisted performance. Gastrointestinal endoscopy possesses a competence-maintenance curriculum adopted by consensus. Horizontal AI governance possesses an obligation to remain aware.
What aeronautics eventually admitted is that control is not external to the system controlled. AI law reasons in terms of a static architecture of responsibilities: it assigns roles and verifies their existence at a given instant. It would need to reason in terms of a closed-loop dynamic system, in which oversight capacity is an endogenous state. This corrects an earlier position: under the name Authority-Influence Gap, I had distinguished formal authority from effective cognitive control as an observed state. It is a trajectory. Authority is stable by juridical construction, the capacity to exercise it follows the dynamics described above, and the gap opens mechanically. As for human oversight, it figures on the compliance balance sheet as a permanent asset, non-amortisable, never subjected to an impairment test, and whose use is exactly what the deployment is intended to reduce.
Regulation (EU) 2026/1744, which entered into force on 27 July 2026, postponed to 2 December 2027 the application of the main obligations for stand-alone high-risk systems under Annex III, without modifying the substance of Article 14. Eighteen additional months of delegation before the asset is inventoried.
One must begin with the question no doctrine of deskilling ever poses and which conditions all the rest. Nothing obliges an organisation to preserve one hundred per cent of its historical autonomous capacity: it no longer knows how to etch its microprocessors, repair its lifts or compute a payroll by hand, and that is perfectly rational. Not every depreciation is a debt. The debt appears only at the crossing of a threshold, C_available < C_required(resilience).
Six parameters conceptually determine its level, none of which is cognitive: criticality of the task, frequency of system failures, detectability of those failures, time necessary for recovery, availability of independent alternatives, reversibility of the damage. What must be defined is therefore not a level of human competence, it is a minimum independent capacity for resilience: the capacity the organisation must retain in order to detect a failure, interrupt the system, maintain a safe state and restore an acceptable service. It is the exact analogue of the minimum viable service in IT resilience, transposed to competence.
This displacement lifts most of the objections addressed to doctrines of competence maintenance. It obliges the retention neither of all former competences nor, in each supervisor, of the capacity to redo everything, since competence can be distributed and escalation organised, and the threshold is proportionate to risk rather than uniform. The objective is not to preserve maximum human autonomy, it is to preserve the minimum autonomy necessary for the mastery of risk. It is distinguished, finally, from the comprehension threshold proposed by Lin and co-authors under the name Capability-Comprehension Gap (arXiv:2602.00854), which is epistemic and bears on the contestability of an output: the resilience threshold is operational and bears on the capacity to hold the system in a safe state without it. Both are necessary and neither substitutes for the other.
First renunciation, the single indicator. The override rate seems the obvious candidate and it is worth nothing on its own: a decline may signify an improvement of the model, a degradation of the controller, excessive trust, an increased workload, an organisational cost of contradiction, an interface that discourages disavowal, a selection of simpler cases, or time pressure. An indicator that admits eight competing explanations is not an indicator. What must be built is a panel, that is to say an observability of human oversight: blinded non-assisted performance, override rate, share of correct overrides, false-disavowal rate, calibration between stated confidence and accuracy, time to detection, capacity for independent justification, and performance after prolonged interruption of assistance, which is the only means of estimating τ_recovery.
Second point, split the protocol in two. Measuring the operator without the system answers the question whether they still know how to do it, which is not the regulatory question. Whether they know how to supervise under real conditions is a different estimand. What is needed is a test of autonomous competence, which measures the stock, and a test of supervisory effectiveness, which injects into a controlled environment a defined proportion of erroneous outputs and measures detection rate, correction rate and time to detection. A form of red teaming applied not to the model but to the supervisor.
Still, the right errors must be injected, failing which the test measures the trivial. A crude hallucination teaches nothing. A false answer, perfectly argued, correctly sourced except on one discreet premise, teaches a great deal. The protocol must stratify, E = {crude, plausible, argued from authority, omission, framing, correlated}, and produce not a score but a probability of detection conditioned on the difficulty and the type of error. An omission and a factual untruth may present the same nominal difficulty and call upon different competences: what one obtains is not a curve but a surface of supervisory performance. One then obtains what compliance lacks today: not the attestation that a human was present, but the characterisation of what they detect and what escapes them.
The aeronautical transposition must not, however, be literal. Requiring a volume of cases handled without assistance would be premature: ten trivial cases are not worth one critical case, and the events one must know how to detect are often too rare to be encountered naturally. What aviation teaches is not a raw quota but a requirement of recurrent proficiency, complemented by simulation when natural exposure is insufficient. Duly noted: synthetic cases, rare scenarios and injected errors become the instrument for maintaining human competence in the face of AI. Simulated populations, long debated as a substitute for data, find here a less contestable function.
Hence a condition that all the preceding reasoning makes obligatory. If the supplier of the system also produces the test cases, the injected errors and the scoring scale, one reconstitutes exactly the coupling one claims to be measuring: the evaluation apparatus becomes an organ of the system evaluated. Independence of the test with respect to the system tested is a design property, not a good practice. It is obtained through distinct models or suppliers, externally validated case sets, an external ground truth, or synthetic generation followed by independent validation. Failing that, one asks the model to write the examination that will verify whether the human can spot its errors, which would constitute a rather creative evaluation protocol.
Finally, one must say what this costs, since that is the point on which a chief financial officer will decide. A minimum independent capacity is not free: cases handled without assistance, periodic simulations, double readings, adversarial exercises, expert time, back-up capacity. Its maintenance constitutes a recurrent resilience expense, and that expense belongs to the total cost of automation, not to the overheads of compliance. It is the same operation described in Delegation is not abdication: what delegation does not cancel must be provisioned somewhere, and where it is booked decides who pays for it.
Four limits, held. The empirical base is thin and divergent: two serious trials conclude in opposite directions, including on experienced populations, and nothing today says which of the two describes the general case. A result obtained on a perceptual act does not transport mechanically to analytical reasoning, and that is an extrapolation named as such. The constants of depreciation and restoration are unknown, the observed horizons short whereas the thesis bears on a cumulative phenomenon. Finally, the position defended is easier to state than to refute, which is exactly the reproach addressed to the existential discourse: I do not claim to escape it, I claim to supply the means of getting out of it.
It remains to arrange the objects in the right order. Extinction, dispossession and cognitive dependence are not three comparable risks: the first is a consequence, the second a state, the third a mechanism, and aligning them would amount to committing the error reproached at the outset. The correct form separates three levels. Mechanisms, among them cognitive depreciation, automation bias and contamination of the supervisory signal. An intermediate state, the loss of effective supervision. Consequences, from undetected errors through to catastrophe. Extinction occupies the last cell of the third column, with an undecidable probability. Cognitive substitution occupies the first cell of the first, with a mechanism documented since 1983, two trials that contradict each other for want of a common protocol, a systematic review that records the heterogeneity of the field, a regulatory precedent in a neighbouring sector, and a professional curriculum that proves the thing is regulable. Working on the first column is not renouncing the third, it is the only way of getting there.
Human oversight is not a guarantee that one enters into the file. It is a state variable of the system it is supposed to control, and the question is not whether it declines. The question is what level must be held, and having an instrument to establish that it is. We shall keep the authority. It is not necessary for the machine to take control if we cease to maintain what would allow us to exercise it.
Doctrinal notes and explorations on AI in regulated systems. Once or twice a month. One-click unsubscribe.