TL;DR: The EU's draft human oversight standard mandates that oversight be verified but leaves testing under realistic conditions optional. I propose four wording changes and explain how to submit them before the comment windows close (some shut within days!). Below: the finding, the fix, how EU standardisation works, and what no standard can fix because Article 14 left it out.
Imagine a triage tool used in a health emergency: it passes every requirement in the current draft, verified the way the draft permits — maybe in a quiet room, several clinicians, a few minutes per patient. The system passes with flying colours and is quickly deployed in the field.
Two paramedics find themselves as the first responders of a mass casualty incident, a collision between a bus and an 18-wheeler truck with 24 injured. The tool tags a 29-year-old with chest pain and stable vitals as "green", but it's wrong. Internal bleeding that the paramedic would have caught had she been in a quiet room gets missed. The paramedic has 30 seconds, a queue of patients and a screen that tells her "low priority". She doesn't overthink it.
While verification is mandatory, testing with representative users under realistic conditions is recommended or noted as possible, not required. Automation bias rises sharply under time pressure and task load (Parasuraman & Manzey 2010[1]), which is precisely the condition the verification never checked.
Nothing malfunctioned. The oversight the standard verified was not the oversight that would play out in real life, at the roadside.
The draft — prEN 18229-3 — is the European standard that will spell out what Article 14's 'human oversight' means in practice (how it got there is explained below).
Notably, JTC 21 (the body in charge of the draft) went further than instructed in the standardisation request. The current draft addresses many of the issues that are not delineated in Article 14, such as having reaction timeframes for intervention as a key element of human oversight, interface design, and deployer-side obligations. These are all important advances, yet there is still a gap.
The document itself defines the convention: "shall" is a requirement, "should" a recommendation, "can" a possibility. It mandates the result (oversight must be verified), but every test that would reveal whether human oversight survives contact with a stressed human — verification under realistic conditions, automation-bias testing, training content — is recommended or noted as possible, not required. In a conformity assessment (usually done in-house, a topic for another piece), testing under real-life conditions is essentially optional. The draft even describes the test itself as an example: realistic conditions, flawed outputs mixed in. It just doesn't require it.
The draft is still in enquiry. Now is the time to comment via your national standards body (the organisation that represents your country in CEN-CENELEC and consolidates comments for JTC 21). This is the only public window before the formal vote.
The CEN-level enquiry closes on 22 October, but national bodies close earlier and dates vary: Netherlands (NEN) 22 September, Ireland (NSAI) 24 September, Sweden (SIS) 7 October, Norway (Standard Norge) 8 October, Spain (UNE) 9 October, France (AFNOR) 12 October. Check your national body's date today and treat it as the real deadline.
The fix is small: change "can/should" to "shall" in 5.3.2.3.2, 5.3.4.5, 5.5.2 and 5.6.2 (the numbers are the draft's own section headings), so that testing with representative users under realistic conditions becomes a requirement.
You can access your national body's enquiry portal via CEN-CENELEC's page here.
National body portals ask for clause, comment and proposed change, one entry per clause. The draft is in English and JTC 21 works in English, so submit your comment in English even on a national portal (no need to translate the quotes). Each block below is one complete entry, paste as is.
Example screenshot of a national body's form filled with one proposed change (in this case it's the Spanish portal, UNE). Yours will look similar:
The legislation-to-standards pipeline follows this structure:
The EU AI Act (Article 14) → C(2025)3871 → prEN 18229-3
The Standardization Request [C(2025)3871]. The European Commission's formal instructions to the relevant bodies for harmonised standards. The relevant bodies are:
These are independent bodies working jointly through CEN-CENELEC JTC 21 (Joint Technical Committee) on the standardisation of, amongst other items, human oversight.
In short, Article 14 requires that AI systems be effectively overseen through their interface, that risks are prevented or minimised and that oversight measures are commensurate with those risks and are provided to the deployer.
The Article has no specifics regarding reaction timeframe, interface requirements or automation bias training.
This can be excused by understanding that such legislation is horizontal in nature — intended to act as a framework within which artificial intelligence systems operate — yet naming these issues would have cost nothing.
From the Commission Implementing Decision[2]:
2.5. Human oversight
The harmonised standards and standardisation deliverables in this area shall set up specifications for human oversight. Those specifications shall comprehensively cover all elements referred to in Article 14 of the Regulation (EU) 2024/1689.
While the mandate says to "comprehensively cover all elements of Article 14", there are no design-in conditions, so the vagueness of the Act is inherited in the mandate (as I previously addressed in this piece).
The European Standard draft on human oversight, prEN 18229-3 (Part 3 of CEN-CENELEC's AI trustworthiness framework) is currently in development, alongside other standard requests under C(2025)3871.
A standard gains legal effect in two steps:
Even with mandatory testing, issues remain. The draft (and the Article it originates from) assumes a designated individual will take responsibility — and, if anything fails, the burden — of working with these systems, as pointed out by Elish (2019)[3]. Humans become "moral crumple zones", soaking the blame for mishaps that actually originated in the development and management structure of the AI system, not in human operators.
Oversight might not fail individually, but institutionally: through problems in staffing, load and authority (as pointed out by Green, 2022)[4].
The Article brings up institutional mechanisms once, in paragraph 5, for high-risk AI systems that use remote biometric identification systems:
[...] no action or decision is taken by the deployer on the basis of the identification resulting from the system unless that identification has been separately verified and confirmed by at least two natural persons with the necessary competence, training and authority.
The requirement for a separate verification by at least two natural persons shall not apply to high-risk AI systems used for the purposes of law enforcement, migration, border control or asylum, where Union or national law considers the application of this requirement to be disproportionate.
This last section adds an institutional mechanism, but immediately makes a carve-out for law enforcement, migration, border control and asylum. It is made optional in precisely the situations that have high-stakes and intense cognitive load, when human oversight is predicted to fail more often (Green 2022[4], Parasuraman & Manzey 2010[1]).
A standard can't fix what the original article waived. For now, we must focus on fixing the four "can/should"s to "shall"s (5.3.2.3.2, 5.3.4.5, 5.5.2 & 5.6.2).
Comment via your national body using the text above in the "Ready-to-paste comment text" section of this piece.
Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human factors, 52(3), 381-410. https://doi.org/10.1177/0018720810376055
Commission Implementing Decision C(2025)3871 on a standardisation request to CEN and CENELEC as regards high-risk AI systems in support of Regulation (EU) 2024/1689 (the AI Act), repealing Implementing Decision C(2023)3215. https://ec.europa.eu/transparency/documents-register/api/files/C(2025)3871_0/de00000001072818?rendition=false
Elish, M. C. (2019). Moral crumple zones: Cautionary tales in human-robot interaction (pre-print). Engaging Science, Technology, and Society (pre-print). https://dx.doi.org/10.2139/ssrn.2757236
Green, B. (2022). The flaws of policies requiring human oversight of government algorithms. Computer Law & Security Review, 45, 105681. https://doi.org/10.1016/j.clsr.2022.105681
CEN-CENELEC (2026). prEN 18229-3: AI trustworthiness framework — Part 3: Human oversight. Draft European Standard, Enquiry stage, CEN-CENELEC JTC 21, July 2026.