A ghost requirement haunts European artificial intelligence regulation: the human oversight mandate that stakeholders invoke constantly but the legislation itself never actually stipulates. Industry players, procurement teams, and compliance presentations all reference this concept. They attribute it to the Act. They cite it in purchasing documents. They pledge it in vendor proposals. They display it on slides with article citations. Yet Article 14 contains no such language. The phrase does not appear once, nor does any variant like 'human on the loop'.
What Article 14 actually specifies is a set of competencies required of whoever performs oversight. That individual must grasp the system's strengths and weaknesses. That individual must stay alert to automation bias—the well-documented inclination to place excessive confidence in machine-generated results. That individual must be capable of understanding system outputs, making choices about whether to apply them, rejecting or modifying them, and halting their execution, including via an emergency stop function. In certain biometric identification scenarios, two distinct natural persons must provide confirmation.
Examine that checklist closely and identify what is absent. Each requirement characterizes a qualified overseer. None of them specifies when the system may proceed. A person can meet every requirement on that list by examining exceptions retroactively, and the Act provides no guidance otherwise.
The actual source of the terminology
During November 2012, Human Rights Watch and Harvard's International Human Rights Clinic released a document advocating for an advance prohibition on fully autonomous weapons systems. The publication, titled Losing Humanity, introduced the vocabulary now employed to reassure hospital leadership.
The report established three categories: Human in the loop describes a machine that identifies targets and applies force exclusively upon human authorization. Human on the loop describes a machine that identifies and applies force under human observation, with the capacity to reverse the action. Human out of the loop describes a machine that performs both functions without human participation.
Human Rights Watch systematized and promoted this framework rather than originating it, since in and out of the loop derive from earlier control-systems vocabulary. The report's contribution was assigning moral significance to the preposition. It also included an observation that has proven disturbingly prescient: an on-the-loop system with merely token oversight functions, practically speaking, as a fully autonomous system.
Thirteen years later, two of those three formulations appear interchangeably in artificial intelligence marketing, deployed by individuals who would be astonished to discover these phrases originally described the circumstances under which a device may take a life.
The distinction between in and on
In the loop, the starting condition is non-action. Activity does not commence until a person authorizes it. The human functions as a checkpoint, and the system's capacity is constrained by the speed of human deliberation.
On the loop, the starting condition is action. The system operates, and the person steps in if complications emerge. The human functions as a restraint, and the system's capacity is unrestricted.
Being in the loop means the system waits for you. Being on the loop means the system has already gone, and you are the reason it might come back.
This distinction constitutes the entire business rationale for the second model, and it also explains why the safeguarding it offers deteriorates so predictably. A checkpoint requires activation, and activation generates documentation. A restraint merely needs to exist, and existence is extremely hard to verify. Token oversight and genuine oversight appear indistinguishable on an organizational diagram.
Why this absence likely was intentional
The legislative authors probably did not overlook this. Mandating human authorization for every high-risk decision would render many of these applications impractical, and an unused application avoids regulatory scrutiny—it simply gets abandoned. Article 14's specifications are deliberately calibrated to the degree of danger and the degree of system independence, representing sensible technical design and simultaneously the precise pathway through which an organization can remain on the loop while satisfying every requirement.
The United States Department of Defense, confronting the identical challenge in the field where this vocabulary originated, took a more aggressive stance and rejected the loop framework altogether. Directive 3000.09 mandates sufficient human judgment regarding the deployment of force. This represents either a straightforward acknowledgment that the binary framework was always inadequate, or a technique for avoiding specification of which category applies.
The conference presentation that made this concrete
The reason for examining a thirteen-year-old weapons document stems from a presentation at this year's European Health Psychology Society conference. Alex Gillespie of the London School of Economics delivered the keynote, titled "Using AI to analyze patient voice and hospital listening: insights into patient safety."
The foundational research dates back nearly a decade and involves no machine learning whatsoever. Gillespie and Tom Reader developed the Healthcare Complaints Analysis Tool, released in BMJ Quality and Safety in 2016, to categorize what patients and families communicate when submitting grievances. Their subsequent investigation spanning 1,110 complaints from 56 NHS trusts demonstrated that complaints identify problems that incident reports overlook, and a further examination across 59 trusts found that the clinical severity described in complaints represented the sole patient-provided metric correlated with hospital-level mortality.
This correlation operates at the hospital level and in cross-sectional form—one stage away from demonstrating complaints forecast mortality. The genuine insight is a safety indicator embedded in complaint correspondence, with insufficient human capacity to process it at volume.
This is where a large language model becomes the natural solution, and where the preposition transforms from a linguistic matter into something consequential.
The evidence that warrants caution
The most compelling findings originate from Gillespie's own research group, which lends them credibility. A group comprising Hannah Bunt, Alex Goddard, Reader and Gillespie assessed GPT-4o performance relative to human annotators across three classification assignments, each containing 1,500 items sourced from NHS complaint records.
In detecting reported speech—whether a section reproduces something someone stated—the model achieved an F1 score of 0.91. In detecting repair, the same. Performance approximated human-level agreement.
In categorizing harm—whether the complaint documents a patient experiencing injury—weighted kappa ranged from 0.49 to 0.57contingent on the prompt formulation, and the model did not merely diverge from human annotators—it systematically inflated harm estimates. Performance was middling at best. Unremarkable for preliminary investigation, and unsuitable for a triage mechanism.
A separate validation study points toward the same conclusion, showing strong alignment on complaint classification and weak alignment on severity and harm assessment. That source should be independently confirmed before relying upon it, though it aligns with the validation research.
The pattern itself is instructive. The model demonstrates dependability where language is unambiguous and unreliability where assessment demands clinical judgment. It excels at the components that matter least and falters at the component the entire initiative exists to accomplish. No one should characterize this AI as reading patient complaints with human-equivalent proficiency, because the research paper conveys nearly the reverse precisely in the domain that matters.
This finding is why the absent phrase becomes significant. A system whose poorest classification capability concerns harm is a system whose human overseer should function as a gate. What the Act permits is a brake.
The second hazard, left unregulated
The other concern surfaced at that presentation had nothing to do with oversight design, and it is the one that has lingered with me.
The hazard is that an academic environment prioritizing publications over practical results will validate these systems using the metrics that generate publications. An F1 of 0.91 on reported-speech classification qualifies as publishable. A weighted kappa of 0.57 on harm becomes a footnote. Both appear in the identical paper, which reflects well on the researchers, and only one will make it into the purchasing presentation.
This represents the publication-versus-outcomes dilemma in its most tangible expression: which of two truthful metrics gets transmitted forward, determined by an incentive system that has no stake in patient safety whatsoever, without any deception or exaggeration necessary.
What becomes of the preposition
It merits asking plainly about any system promoting human oversight: does the system pause for approval, or does it proceed? If it proceeds, who is monitoring, how many outputs per hour, and what consequences face that person if they object too frequently?
Article 14 will not pose those questions on your behalf. It characterizes a qualified overseer and leaves the operational model to the purchasing organization.
What lingers is that a distinction crafted to regulate lethal weapons now quietly governs how a complaint from a grieving family gets processed, that the 2012 caution regarding token oversight accompanied it unchanged, and that the terminology everyone invokes to demonstrate human authority turns out to be absent from the legislation they are citing.
Source: Silicon Canals



