TL;DR: In this piece we examine how human oversight translates from a vague mandate to specific standards in its current draft form, what is promising and a call to action before the 22 October deadline to make testing of human oversight effectiveness in AI systems obligatory. We explore the inner workings of EU Standardisation from from the EU AI Act (Article 14) to the standardization request, and finally the current draft. Lastly, thoughts on institutional mechanisms (or lack thereof) from the Article itself, and what can be done.
Imagine a triage tool used in a health emergency: it passes all the necessary requirements in the current draft, verified the way the draft permits; maybe in a quiet room, with multiple clinicians and a few minutes per patient. The system passes with flying colours and is quickly deployed in the field.
Two paramedics find themselves as the first responders of a mass casualty incident, a collision between a bus and an 18-wheeler truck with 24 injured. The tool tags a 29 year old with chest pain and stable vitals as "green", but it's wrong. Internal bleeding that the paramedic would have caught had she been in a quiet room gets missed. The paramedic has 30 seconds, a queue of patients and a screen that tells her "low priority". She doesn't overthink it.
None of this was tested under real-life conditions. While verification is mandatory, testing with representative users under realistic conditions is noted as possible, not required. Automation bias rises sharply under time pressure and task load (Parasuraman & Manzey 2010), which is precisely the condition the verification never checked.
Nothing malfunctioned. The oversight the standard verified was not the oversight that would play out in real life, at the roadside.
The draft — prEN 18229-3 — is the European standard that will spell out what Article 14's 'human oversight' means in practice (how it got there is explained below).
Notably, JTC 21 (the body in charge of the draft) went further than instructed in the standardization request. The current draft addresses many of the issues that are not delineated in Article 14, such as having reaction timeframes for intervention as a key element of human oversight, interface design, and deployer-side obligations. These are all important advances, yet there is still a gap.
The document itself defines the convention: "shall" is a requirement, "should" indicates a recommendation and "can" indicates possibility or capability as can be read in clause 0.2 (the numbers throughout are the draft's own section headings). It mandates the result (5.3.2.3, 5.3.4.4), but every test that would reveal whether human oversight survives contact with a stressed human — realistic-conditions verification (5.3.2.3.2, 5.3.4.5) and automation bias testing (5.5.2) — is noted as possible, not required. In a conformity assessment (done in-house, a topic for another piece), testing under real-life conditions is essentially optional.
The draft is still in enquiry until 22 October. Now is the time to comment via your national standards body, which consolidates comments for JTC 21. This is the only public window before the formal vote.
The fix is small: change "can" to "shall" in 5.3.2.3.2, 5.3.4.5 and 5.5.2, so that testing with representative users under realistic conditions becomes a requirement.
You can access your national body's enquiry portal via CEN-CENELEC's page here.
Comment on prEN 18229-3 (enquiry, closes 22 Oct 2026): — 5.3.2.3.2: "can be performed" → "shall include" testing with representative designated persons under realistic oversight conditions. — 5.3.4.5: "can be conducted" / "should be tested" → "shall". — 5.5.2: "can conduct" automation-bias testing → "shall". — 5.6.2 (consider): "should include" training content → "shall". Rationale: the draft mandates verification of outcomes (5.3.2.3, 5.3.4.4) but leaves the only method that exposes failure under time pressure and task load optional. Automation bias rises sharply under exactly these conditions (Parasuraman & Manzey 2010).
The legislation-to-standards pipeline follows this structure:
The EU AI Act (Article 14) → C(2025)3871 → prEN 18229-3
The Act is the legislation that was adopted by the European Parliament and the Council, which is difficult to change once passed.
Next comes the Standardization Request [C(2025)3871], which is the European Commission's formal instructions to the relevant bodies for harmonized standards. These will be used across the European Union.
The relevant bodies are:
These are independent bodies working jointly through CEN-CENELEC JTC 21 (Joint Technical Committee) on the standardization of, amongst other items, human oversight.
Lastly, the request is drafted into a European Standard Project or prEN. In this case, the project regarding human oversight is prEN 18229-3.
One last important distinction regarding legislation vs. standards:
What Article 14 states regarding human oversight (paraphrased for shorter length); it asks for AI systems with an interface that can be used effectively to oversee AI systems, prevent risks, have oversight measures commensurate with risks and that are provided to the deployer.
The Article has no specifics regarding reaction timeframe, interface requirements or automation bias training, just to name a few issues that can disrupt functional human oversight. This can be excused by understanding that such legislation is horizontal in nature, intended to act as a framework from which artificial intelligence systems operate. This means it does not need to go into specifics, yet they could still have been named in the legislation to make sure they are not ignored.
From the Commission Implementing Decision[1]:
2.5. Human oversight
The harmonised standards and standardisation deliverables in this area shall set up specifications for human oversight. Those specifications shall comprehensively cover all elements referred to in Article 14 of the Regulation (EU) 2024/1689.
While the mandate says to "comprehensively cover all elements of Article 14", there are no design-in conditions, which means the vagueness of the Act is inherited in the mandate. This language is especially worrying given the lack of explicit language in Article 14 to address the problems with human oversight, as I previously addressed in this piece here.
The European Standard draft on human oversight, prEN 18229-3, (Part 3 of CEN-CENELEC's AI trustworthiness framework) is currently in development, alongside other standard requests under C(2025)3871.
A standard gains legal effect in two steps:
Even with mandatory testing, issues remain. The draft (and the Article it originates from) assumes a designated individual will take responsibility — and, if anything fails, the burden — of working with these systems, as pointed out by Elish (2019)[2]. Humans become "Moral Crumple Zones", soaking the blame for mishaps that actually originated in the development and management structure of the AI system, not in human operators.
Oversight might not fail individually, but institutionally: through problems in staffing, load and authority (as pointed out by Green, 2022)[3].
The Article brings up institutional mechanisms once in the 5th point (the one we didn't mention earlier):
5. For high-risk AI systems that use remote biometric identification systems there are stronger rules:
[...] no action or decision is taken by the deployer on the basis of the identification resulting from the system unless that identification has been separately verified and confirmed by at least two natural persons with the necessary competence, training and authority.
This rule does not apply to high-risk AI systems used for the purposes of law enforcement, migration, border control or asylum, where the union or national law considers the application of this requirement to be disproportionate.
This last section adds an institutional mechanism, but immediately makes a carve-out for law enforcement, migration, border control and asylum — it is made optional in precisely the situations that have high-stakes and intense cognitive load, when human oversight is predicted to fail more often (Green 2022[3], Parasuraman & Manzey 2010[4]).
A standard can't fix what the original article waived. For now, we must focus on fixing the three 'can's (5.3.2.3.2, 5.3.4.5, 5.5.2) before the 22nd of October. Comment via your national body using the text above.
Elish, M. C. (2019). Moral crumple zones: Cautionary tales in human-robot interaction (pre-print). Engaging Science, Technology, and Society (pre-print). https://dx.doi.org/10.2139/ssrn.2757236
Green, B. (2022). The flaws of policies requiring human oversight of government algorithms. Computer Law & Security Review, 45, 105681. https://doi.org/10.1016/j.clsr.2022.105681
Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human factors, 52(3), 381-410. https://doi.org/10.1177/0018720810376055
CEN-CENELEC (2025). prEN 18229-3: AI trustworthiness framework — Part 3: Human oversight. Draft European Standard, Enquiry stage (Stage 40), CEN-CENELEC JTC 21. Enquiry opened 30 July 2025, closes 22 October 2026. Brussels: CEN-CENELEC Management Centre.