Article — Position paper · ○ Open access

Attributing a decision is not organizing its oversight

What a human-oversight guarantee must specify, and what the signatory's name leaves undetermined

Jérôme Vetillard · · Twingital Institute · 6 pages · 11 min read
🇫🇷 Lire en français ↓ Download PDF

Abstract

An agreement between the State of Utah and the company Doctronic authorizes an artificial intelligence system to renew prescriptions, and sums up its safeguard in one line: a licensed physician signs. The same agreement nevertheless describes three very different ways of organizing oversight, full prior review, full retrospective review, then monthly sampling, while the name attached to the prescription never changes.

The thesis fits in one sentence. Attributing a decision to a signatory tells you nothing about how its oversight is organized. A human-oversight guarantee becomes usable, in a contract as in an audit, only when it states the timing of the intervention, its coverage and the reviewer’s capacity to act, complemented by the quality of the review and by measurable purposes.

The Doctronic pilot: a one-line guarantee, a three-regime agreement

On 6 January 2026, the Utah Department of Commerce made public a regulatory mitigation agreement with Doctronic. Its scope is narrow: 30-, 60- or 90-day renewals of medications already prescribed, for chronic conditions and mental health, within an approved formulary, excluding controlled substances, new prescriptions and any change to the treatment plan.

The program’s public page states the guarantee this way: renewals are signed and approved by a licensed physician, “either directly or vicariously through the AI system’s protocol”. The contract is more precise than the slogan. It chains three review regimes. First, a physician reviews every decision before it is transmitted to the pharmacy. Second, every decision is still reviewed, but after the fact. Third, review covers a monthly sample of 5 to 10 % of renewals, backed by a quarterly analysis of escalated cases and an annual review of the indicators.

Two status caveats frame the reading. The progression thresholds between regimes, initially expressed in numbers of patients, have since been amended, and the official page now describes eligibility by medication group: the analysis concerns the review architecture, not the thresholds in force today. And as of 15 September 2026, no public source consulted establishes that any change of regime has occurred. These are planned regimes, not observed practices.

What moves while the signatory’s name stays put

From the first regime to the second, coverage remains total and only timing changes: review moves from before dispensing to after. For the file under review, it stops being a barrier and becomes a finding. From the second to the third, two variables shift together. Coverage drops from every decision to a sample, and the scheduled cadence becomes monthly, which delays detection compared with a review conducted shortly after issuance, unless another channel catches the error first.

Those other channels exist and persist across all three regimes. The pharmacist can refer a case to a physician, the patient can request a review at any time, and the system automatically escalates cases whose clinical complexity exceeds predefined thresholds. The 5 to 10 % figure therefore describes one modality of scheduled review, not the whole of human intervention. Reading it requires its denominator, and adding up the channels requires not counting the same decision twice.

Throughout, one term stays fixed: the physician’s name on the prescription. Whether the obligations and liability attached to that name remain identical from one regime to the next is a separate question, and the agreement does not settle it. Hence the central proposition: two systems can share the same attribution, and even the same review coverage, without offering the same preventive capacity. Attributing a decision and organizing its oversight are two different objects, and the first says nothing about the second. It is the mirror image of the problem described in The agent that decides has no signatory: there, a decision was looking for its signatory; here, a signatory does not say how the decision was reviewed.

What must a human-oversight guarantee specify?

Saying that a system does or does not include human oversight applies a binary vocabulary to an object that is not binary. Describing how oversight is organized takes at least three dimensions, plus two characteristics they do not capture.

ElementWhat must be declaredQuestion it answers
Timingposition of the intervention relative to the person’s exposure, associated delayscan oversight still prevent the effect?
Coveragedecisions reviewed, sampling method, complementary channelswhat share of decisions does it reach, and by what rule?
Capacity to actauthority to impose a correction, technical means, execution and verification of the correctionis a detected error actually corrected?
Quality of reviewcompetence, available information, effective detection ability, independence of judgmentwhat does the reviewer actually catch?
Purposes and measurementobjectives pursued, indicators of their achievementwhat is being protected, and how do we know?

Review quality cannot be separated from the three dimensions: constrained review time reduces its depth, and high coverage increases its load. Yet it does not follow from them, since two systems identical in timing, coverage and means of action can detect very different proportions of errors. That is the ground covered by Human oversight is not a presence, which treats the sensitivity of oversight as a quantity to measure, and by Human oversight is not external to the system, which shows that this capacity depreciates with use. The same pilot had already served, in Outside doctrine, before doctrine, to describe oversight that is technically present and substantively absent, the clinician validating an output whose genesis is no longer legible to them. That piece dealt with what the reviewer can see; this one deals with where, when and over what the reviewer has been placed.

This grid is a specification floor, not a complete characterization of the guarantee. It already demands more than a name does. Its domain of validity is explicit: it concerns what an attribution can establish, not the safety of the Utah pilot. Nor does it disqualify supervision by sampling. Exception-based monitoring equipped with a capacity to stop the system can be substantial, and substantial oversight does not require universal review.

Ex ante and ex post control: an old distinction

The principle is not new. French public accounting separates ex ante from ex post financial control. Internal audit distinguishes preventive from detective controls. Financial markets have long contrasted pre-trade with post-trade compliance, precisely because the same coverage rate does not protect in the same way on either side of execution. On this very pilot, Michelle M. Mello has already named the shift from review before action to retrospective review, and noted that no independent evaluation is planned after deployment (JAMA Health Forum, 19 March 2026).

Medicine also knew arrangements in which a professional answers for acts they do not review one by one, such as standing orders or collaborative practice protocols. The delegate there remains a professional who can refuse, and whose own appraisal is itself a review. What automation widens is therefore not the option of placing human intervention at one point of the circuit or another, which was already an organizational choice. It is the ability to move that point, or to remove individual interventions, without the circuit putting up any resistance. Human organization made the control point costly to move; software makes it a setting. The mechanism echoes Friction was the guarantee.

Specify, measure, evaluate: three operations the debate conflates

The strongest objection is clinical. What matters is not how many files a physician reads but the net balance: errors avoided or introduced, quality of referral, treatment interruptions spared, burden shifted onto pharmacists and patients. A count of files reviewed is a process measure, not a measure of clinical benefit.

The objection is sound, and it narrows the reach of the argument. It does not make the net balance sovereign, though: a favorable average net benefit can conceal harm concentrated in a subgroup, or degraded access to redress.

The debate clears up once three operations are separated: specifying the planned oversight, measuring its actual execution, evaluating its consequences. The grid above serves the first two. It does not make the clinical guarantee assessable before any data exist, and a delay declared in a contract does not establish an observed delay. Conversely, one can measure practices, delays and errors without prior specification; what is then missing is a declared reference against which to judge whether practice keeps the promise. An unspecified guarantee is not false: it is unverifiable. This is the logic of Evidence is a conditional promise, applied to oversight.

Whom does a late review protect?

Correcting the file under review and preventing errors in subsequent files are not the same operation. A late review loses its interception capacity on the case it examines. It keeps, and sometimes increases, another value: limiting consequences that are still unfolding, detecting a systematic rather than an isolated error, triggering suspension of the system, preventing errors to come. Monthly sampling is not a prior barrier for every decision; it can be a good instrument for monitoring the system.

Two populations are thus protected by different mechanisms, the patient whose file is read and the patients who come next, while the system itself is the object being monitored. A single oversight arrangement can serve several of these purposes at once. A useful specification does not choose between them: it declares which ones it pursues and how their achievement is measured. The capacity to suspend, in particular, is worth something only if it has been tested, because you govern only what you can still redirect.

AI Act Article 14: a proportionate requirement, a single numbered rule

Article 14 of Regulation (EU) 2024/1689 requires high-risk AI systems to be capable of being effectively overseen by natural persons during the period in which they are in use. Paragraph 3 requires oversight measures “commensurate with the risks, level of autonomy and context of use”. Paragraph 4 lists capacities: to understand the system and monitor its operation, to remain aware of automation bias, to interpret its output, to decide not to use it, to disregard, override or reverse its output, and finally to intervene or interrupt the system through a stop button or a similar procedure that brings it to a halt in a safe state. That last item is the regulatory form of the suspension capacity discussed above: the text requires it as a possibility, and says neither who activates it nor on what signal.

The absence of a universal rate is not the absence of an operational requirement. It hands the specification to the provider and the deployer, as a function of risk and context. Paragraph 5 shows, by contrast, what a complete prescription looks like: for the remote biometric identification systems referred to in point 1(a) of Annex III, the deployer may not act or decide on the basis of an identification unless it has been separately verified and confirmed by at least two natural persons with the necessary competence, training and authority. Timing, coverage, redundancy and even the reviewer’s qualification are all fixed, with an exception where Union or national law deems that double verification disproportionate in certain law enforcement, migration, border control or asylum uses.

Everywhere else, translating the general obligation into verifiable modalities falls to the provider and the deployer. The grid proposed here offers a partial translation, focused on how the intervention is organized, and does not exhaust the regulation. A note on timing: Regulation (EU) 2026/1744, the Digital Omnibus on AI, in force since 27 July 2026, does not rewrite Article 14. It postpones application of the high-risk requirements to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems, and, through a new Article 2(13), opens the possibility of limiting specific requirements of Articles 9 to 15 and 17 to 25 for systems covered by the Union harmonisation legislation listed in Annex I, Section A, where that legislation provides an equivalent or higher level of protection and without reducing the overall level of protection. That limitation requires delegated acts, which the Commission must adopt by 2 August 2027; until then, Article 14 remains fully applicable to those systems from their date of application.

The contractual consequence: a human-oversight clause must spell out its modalities

A contract clause, an attestation or a tender response that promises human oversight becomes usable when it declares its timing, coverage, capacity to act, review-quality conditions and purposes, instead of settling for a designated responsible person. Otherwise, the promise does not allow any preventive capacity to be appraised. An audit can always find the specification insufficient, look for actual practices, examine other dimensions; it will do so without a declared point of comparison. Designating a responsible person settles attribution, not oversight, and delegation is not abdication.

The file supports a single general rule, and it is conditional: the more irreversible the consequences of an error, the more decisive prevention before exposure becomes, and the less retrospective monitoring can stand in for it.

Limits and domain of validity

The file documents what a contract provides and what an operator declares. It documents neither actual execution, nor what supervisory authorities verify, nor what courts sanction. Capacity to act is posited as a dimension, not measured: establishing it would require the actual delays of transmission, dispensing and first dose, the reviewer’s effective authority, their means and the verification of the corrections decided, none of which appears in the sources consulted. Finally, the initial agreement lists its performance indicators, physician-system concordance, approval rate, patient and pharmacist escalation rates, without attaching any numerical threshold to them. That is an observation about that document in its initial version, not a conclusion about the protocol or the criteria in force.

Eleven of the fourteen members of the Utah Medical Licensing Board called for the pilot to be suspended, citing clinical reassessment, the risk of drug interactions and the lack of prior consultation of the Board. The State’s response, dated 21 April 2026, spells out the review regimes and the rate of the first one. Timing and coverage were therefore available as soon as the guarantee came under discussion. It is the signatory’s name that does not carry them: the timestamp on a signature does not, on its own, establish when or how the file was reviewed.

Frequently asked questions

Does a signing physician guarantee effective human oversight?

No. A signature establishes attribution. It does not say whether review happens before or after the patient’s exposure, what share of decisions is reviewed, or whether the reviewer has the authority and means to impose and then verify a correction.

Is sample-based human oversight compatible with Article 14 of the EU AI Act?

Article 14 sets no review rate, except for the double verification of remote biometric identification. It requires measures commensurate with risk, level of autonomy and context of use. Sampling can meet that requirement if it is specified, backed by a capacity to stop the system and complemented by escalation channels; it is not, however, a prior barrier for each decision.

What is the difference between ex ante and ex post review of an automated decision?

Ex ante review can intercept an error before it reaches the person. Ex post review can no longer do so for the case reviewed, but it can limit consequences still unfolding, reveal a systematic error and trigger suspension of the system.

What should a human-oversight contract clause contain?

At a minimum: the timing of the intervention and its delays, the coverage and selection method of reviewed decisions, the authority and means of correction, the conditions of review quality, and the purposes pursued with the indicators of their achievement.

Primary sources

Read the document