Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- EU AI Act Conformity Testing: What Article 9 Requires
EU AI Act Conformity Testing: What Article 9 Requires
The Digital Omnibus moved the high-risk dates. See what Article 9 requires you to test, what prior defined metrics mean, and what a notified body can demand.
Published on:
If your compliance plan says high-risk obligations bite on 2 August 2026, it is working from text that has been replaced. Regulation (EU) 2026/1744, the Digital Omnibus on AI of 8 July 2026, rewrote Article 113(c).
Annex III high-risk systems now fall under 2 December 2027, and Annex I high-risk systems under 2 August 2028. A great deal of published guidance still carries the superseded dates.
That extra runway is not idle time. The conformity machinery applies before the obligations it assesses, and the evidence a QA function has to produce takes longer to build than the documents that describe it.
TL;DR
EU AI Act conformity testing is the work of turning Chapter III obligations into test design and retained evidence for a high-risk AI system. Article 9 is where most of that work lands, because it is the provision that says testing shall happen and sets the conditions it happens under.
- Did the high-risk deadline move?: Yes. Regulation (EU) 2026/1744 replaced Article 113(c), deferring Chapter III Sections 1 to 3 to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems.
- Does the deferral cover all of Chapter III?: No. Section 4 on notified bodies has applied since 2 August 2025, and Section 5 on standards, conformity assessment and registration applies from 2 August 2026.
- Is testing optional?: No. Article 9(6) states that high-risk AI systems shall be tested, and that testing shall ensure they perform consistently for their intended purpose.
- Can the threshold be set after the run?: No. Article 9(8) requires testing against prior defined metrics and probabilistic thresholds appropriate to the intended purpose, and the word prior is doing the work.
- Does the Act name a minimum accuracy figure?: No. It requires metrics and thresholds appropriate to the intended purpose, so any specific statutory accuracy percentage you see quoted is not in the regulation.
- Can an assessor re-run your tests?: Yes. Under Annex VII a notified body that is not satisfied with the provider’s tests may carry out adequate tests itself, with access to the training, validation and testing data sets where necessary.
Write the acceptance criteria before the evidence exists, retain what a third party would need to reproduce the run, and keep the post-market data flowing back into risk analysis.
When Do the High-Risk Rules Actually Apply?
Article 113 provides that Regulation (EU) 2024/1689 as consolidated “shall apply from 2 August 2026”, and the European Commission’s application timeline gives the entry into force date as 1 August 2024. Regulation (EU) 2026/1744, the Digital Omnibus on AI of 8 July 2026, replaced the original point (c), so 2 August 2026 is no longer the application date for Chapter III, Sections 1, 2 and 3.
The amended Article 113(c) splits those sections, Article 6(5) excepted, by classification route. The same conformity-assessment shape under a different regulation is walked through in European Accessibility Act compliance.
- 2 December 2027 - Article 113, point (c)(i) applies those sections “as regards AI systems classified as high-risk pursuant to Article 6(2) and Annex III”.
- 2 August 2028 - point (c)(ii) applies them “as regards AI systems classified as high-risk pursuant to Article 6(1) and Annex I”.
- The text those dates replaced - the original point (c) read “Article 6(1) and the corresponding obligations in this Regulation shall apply from 2 August 2027”, and Annex III systems fell under the general 2 August 2026 date. The dates above are those of the amended text.
Sections 4 and 5 never moved. Planners who read the headline as a wholesale move of Chapter III give themselves time Article 113(c) does not grant.
- Section 4, notifying authorities and notified bodies - has applied since 2 August 2025 under Article 113(b): “Chapter III Section 4, Chapter V, Chapter VII and Chapter XII and Article 78 shall apply from 2 August 2025, with the exception of Article 101”. The Omnibus left that point untouched.
- Section 5, standards, conformity assessment, certificates and registration - the amended Article 113(c) does not reach it, so it applies from the general date of 2 August 2026.
Outside Chapter III, the amended Article 113 resets dates still taken from the 2024 text.
- Chapters I and II - apply from 2 February 2025 under the amended Article 113(a), “with the exception of Article 5(1), first subparagraph, points (ba) and (bb), and Article 5(1a) and (1b) which shall apply from 2 December 2026”.
- That exception - points (ba) and (bb) are prohibitions added by the Omnibus, scoped by Article 5(1a) and (1b).
- Articles 102 to 110 - a new Article 113, third paragraph, point (d), inserted by Regulation (EU) 2026/1744 and absent from the text published in OJ L 1689 of 12.7.2024, provides that they “shall apply from 27 July 2026”.
For a QA function, Annex VII is where those dates turn into test work. Point 4.4 provides that in examining the technical documentation the notified body “may require that the provider supply further evidence or carry out further tests”, and that it “shall itself directly carry out adequate tests, as appropriate” where it is not satisfied with the provider’s tests.
- Annex VII, point 4.3 - “Where relevant, and limited to what is necessary to fulfil its tasks”, the notified body “shall be granted full access to the training, validation, and testing data sets used”.
- Annex VII, point 4.5 - access to “the training and trained models of the AI system, including its relevant parameters” follows a reasoned request, and only “after all other reasonable means to verify conformity have been exhausted and have proven to be insufficient”.
- Article 9(8) - testing “shall be carried out against prior defined metrics and probabilistic thresholds that are appropriate to the intended purpose of the high-risk AI system”.
Metrics and thresholds of that kind are defined before a test run, which puts the work inside your current planning horizon.
Note: This article summarises published provisions of Regulation (EU) 2024/1689 as consolidated, so a QA function can translate them into test design. It is not legal advice. Whether a given system is high-risk, and which obligations attach to your organisation as provider, deployer or importer, is a determination for your own legal counsel.
Article 9 as a Test Specification
Article 9 of the Regulation as published sits in Chapter III, Section 2 and reads as a test specification. Article 9(1) requires that “A risk management system shall be established, implemented, documented and maintained in relation to high-risk AI systems.”
- Article 9(1) and 9(2) - “documented and maintained” lets an assessor ask for the artefact and its revision history. Article 9(2) sets “a continuous iterative process planned and run throughout the entire lifecycle”, “requiring regular systematic review and updating”, a different cadence from a release-only refresh. Point (c) feeds it post-market monitoring data under Article 72.
- Article 9(5), first subparagraph - measures adopted under Article 9(2), point (d), shall be such that “the relevant residual risk associated with each hazard, as well as the overall residual risk of the high-risk AI systems is judged to be acceptable”. Acceptability is judged per hazard and again overall, so one blended pass rate leaves the per-hazard limb unanswered.
- Article 9(6) - “High-risk AI systems shall be tested for the purpose of identifying the most appropriate and targeted risk management measures. Testing shall ensure that high-risk AI systems perform consistently for their intended purpose and that they are in compliance with the requirements set out in this Section.” A run informing no measure sits outside that purpose.
- Article 9(7) - “Testing procedures may include testing in real-world conditions in accordance with Article 60.” The verb is “may”, so the campaign is an option you justify or decline. Article 3(57) defines it as temporary testing of an AI system for its intended purpose outside a laboratory or otherwise simulated environment, which does not count as placing on the market or putting into service where the Article 57 or 60 conditions are met.
- Article 9(8) - testing “shall be performed, as appropriate, at any time throughout the development process, and, in any event, prior to their being placed on the market or put into service”, against “prior defined metrics and probabilistic thresholds” suited to the intended purpose. The “as appropriate” qualifier covers development-phase testing; pre-market testing is “in any event”.
The phrase to convert into a measurable claim is “perform consistently for their intended purpose”. Article 9(6) sets that standard and Article 9(8) the evidentiary form, naming no metric, threshold value or accuracy figure. What follows is test-design practice, not a reading of the Regulation.
- Write the intended purpose down first - consistency is measured against that statement across input slices, model versions and elapsed time.
- Record the acceptability judgement - Article 9(5) calls for a judgement. Naming an owner and a date beside the figure is a practice choice, not something the paragraph specifies.
Article 9(1) and 9(2) come from the consolidated text in force from 27 July 2026, after Regulation (EU) 2026/1744 of 8 July 2026, which leaves 9(6), 9(7) and the Article 3(57) definition unchanged. Article 9(2), point (c), 9(5) and 9(8) follow Regulation (EU) 2024/1689 as published on 12 July 2024.
As amended by that Regulation, Article 113(c) applies Chapter III, Sections 1, 2 and 3, with the exception of Article 6(5), from “2 December 2027 as regards AI systems classified as high-risk pursuant to Article 6(2) and Annex III” and “2 August 2028 as regards AI systems classified as high-risk pursuant to Article 6(1) and Annex I”. The original point (c) said 2 August 2027. Your compliance function owns the legal reading; this is engineering commentary, not legal advice.
Prior Defined Metrics and Probabilistic Thresholds
Article 9(8) of Regulation (EU) 2024/1689 sets a deadline and an evidentiary standard in the same paragraph.
- The deadline - testing of high-risk AI systems “shall be performed, as appropriate, at any time throughout the development process, and, in any event, prior to their being placed on the market or put into service”.
- The standard - “Testing shall be carried out against prior defined metrics and probabilistic thresholds that are appropriate to the intended purpose of the high-risk AI system.”
The word “prior” fixes the order of the metric and the evidence in time. Read at face value, that ordering puts familiar evaluation moves outside what the paragraph describes.
- Setting the bar after the run - reading the evaluation output first, then deciding what counts as acceptable.
- Moving the bar - relaxing a threshold once results land under it, and reporting against the relaxed one.
- Metric shopping - carrying several measures through the run and publishing whichever looked best on this build.
Article 9(8) puts the definition step first, which has a practical effect on how a QA function keeps its records. None of what follows is spelled out in the paragraph; each is what teams do to show the ordering it assumes.
- Dated acceptance criteria - written, reviewed and versioned before the evidence exists, so the date carries as much weight as the numbers.
- Threshold changes - worth logging, because a comparison across runs means something only while the bar held.
- Probabilistic thresholds - the phrase is the regulation’s own, though the Act does not define it; read plainly it points at a bar expressed over a distribution of outcomes.
- Appropriate to the intended purpose - the same figure can be defensible for one deployment and indefensible for another, so the justification is worth keeping beside the value.
- The intended purpose - Article 9(8) measures the threshold against the intended purpose of the high-risk AI system, so whatever a team treats as that purpose is what the number is read against.
Where the Annex VII conformity assessment procedure applies, point 4.4 is where that paperwork meets an outside reader.
- Further evidence - “In examining the technical documentation”, the notified body “may require that the provider supply further evidence or carry out further tests so as to enable a proper assessment of the conformity of the AI system with the requirements set out in Chapter III, Section 2”.
- Its own tests - where the notified body “is not satisfied with the tests carried out by the provider”, it “shall itself directly carry out adequate tests, as appropriate”.
The Act sets no universal accuracy figure for high-risk AI systems, so a statutory accuracy percentage quoted at you is describing something the regulation does not contain.
- The requirement - testing carried out against “prior defined metrics and probabilistic thresholds that are appropriate to the intended purpose of the high-risk AI system”.
- The number - Article 9(8) names none.
The wording quoted here is Article 9(8) as published in OJ L series 2024/1689. Article 113(c), as amended by Regulation (EU) 2026/1744 of 8 July 2026, sets the application dates for Chapter III, Sections 1, 2 and 3, “with the exception of Article 6(5)”.
- 2 December 2027 - AI systems classified as high-risk pursuant to Article 6(2) and Annex III.
- 2 August 2028 - AI systems classified as high-risk pursuant to Article 6(1) and Annex I.
This section describes what the cited text says. It is not legal advice.
Notified Body Access Rights
Annex VII to Regulation (EU) 2024/1689 gives a notified body wider access than most release processes serve. Point 4.3 states that “The technical documentation shall be examined by the notified body” and that, on the conditions the point sets, the body “shall be granted full access to the training, validation, and testing data sets used”. That access may be delivered remotely: point 4.3 continues “including, where appropriate and subject to security safeguards, through API or other relevant technical means and tools enabling remote access”.
- Point 4.3, data set access - the grant is bounded by “Where relevant, and limited to what is necessary to fulfil its tasks”, so its scope is settled in the review.
- Point 4.4, further evidence - the notified body “may require that the provider supply further evidence or carry out further tests so as to enable a proper assessment of the conformity of the AI system with the requirements set out in Chapter III, Section 2”. “May” is permissive, so that request is the body’s call.
- Point 4.4, independent testing - “Where the notified body is not satisfied with the tests carried out by the provider, the notified body shall itself directly carry out adequate tests, as appropriate”. Dissatisfaction switches the text from “may” to “shall”, qualified by “as appropriate”.
- Point 4.5, models and parameters - where such access is “necessary to assess the conformity of the high-risk AI system with the requirements set out in Chapter III, Section 2”, and only “after all other reasonable means to verify conformity have been exhausted and have proven to be insufficient, and upon a reasoned request”, the notified body “shall also be granted access to the training and trained models of the AI system, including its relevant parameters”, subject to “existing Union law on the protection of intellectual property and trade secrets”.
- When this applies - under Article 113(b), the one application-date point the Digital Omnibus on AI (Regulation (EU) 2026/1744 of 8 July 2026) left untouched, Chapter III Section 4 on notifying authorities and notified bodies has applied since 2 August 2025. Chapter III Section 5, on conformity assessment, applies from the general date of 2 August 2026. Article 113(c) as amended defers the Chapter III, Section 2 requirements point 4.4 has the body assess against: 2 December 2027 for systems classified as high-risk under Article 6(2) and Annex III, and 2 August 2028 under Article 6(1) and Annex I.
The text sets no requirement about how you keep any of this, and none of the points quoted here sets a retention period. As an engineering matter, though, a body exercising point 4.4 can only re-run tests against material that still exists and still executes, which is a heavier ask than a legible report. Put these questions to your team.
- Version-pinned data - can you produce the exact training, validation and testing data sets used for one release?
- Reasoned request handling - who receives a reasoned request for model access, and who decides, with your legal counsel, what the point 4.5 intellectual property and trade secret protections cover?
- Retention horizon - how long do those data sets, model artefacts and test environments outlive the release?
Treat this as a timed drill: take a release from six months ago and time how long your team needs to reconstitute its data sets, fixtures and one working test run.
The Post-Market Feedback Loop
Article 9(1) requires a risk management system “established, implemented, documented and maintained in relation to high-risk AI systems”. The engineering side of that feedback path is covered in AI observability. Point (c) of Article 9(2) applies after deployment, requiring “the evaluation of other risks possibly arising, based on the analysis of data gathered from the post-market monitoring system referred to in Article 72”. Your production signal is an input to that process, alongside the pre-market test evidence.
- A dated written evaluation - in practice the step is worked as a dated record naming the risk, the production evidence behind it, and the conclusion reached.
- Usage outside points (a) and (b) - point (a) is the identification and analysis of “the known and the reasonably foreseeable risks” to “health, safety or fundamental rights” under the intended purpose, and point (b) is the estimation and evaluation of “the risks that may emerge” under that purpose “and under conditions of reasonably foreseeable misuse”. Production behavior outside both framings goes back into analysis.
- Movement against the declared thresholds - Article 9(8) requires testing against prior defined metrics and probabilistic thresholds appropriate to the intended purpose of the high-risk AI system, so instrumenting production on those metrics keeps the feed and the test evidence in one vocabulary.
- Residual risk acceptability - Article 9(5), first subparagraph, requires the measures referred to in paragraph 2, point (d), to be such that “the relevant residual risk associated with each hazard, as well as the overall residual risk of the high-risk AI systems is judged to be acceptable”.
- The recorded decision - point (d) is “the adoption of appropriate and targeted risk management measures designed to address the risks identified pursuant to point (a)”. In practice an evaluation closes with a measure, a re-test, or a recorded reason for changing nothing.
Article 9(2) describes the risk management system as “a continuous iterative process planned and run throughout the entire lifecycle of a high-risk AI system, requiring regular systematic review and updating”.
- Give the review an owner and a cadence - Article 9(2) sets review and updating that is regular and systematic, and a review that happens when somebody remembers is hard to describe that way.
- File the analysis with the risk documentation - Article 9(1) requires the risk management system to be documented and maintained, so the point (c) output sits beside the pre-market test evidence.
- Write for an Annex VII examination - where that conformity assessment procedure applies, the notified body “may require that the provider supply further evidence or carry out further tests so as to enable a proper assessment of the conformity of the AI system with the requirements set out in Chapter III, Section 2”.
- Re-test when the evaluation surfaces something - Article 9(6) provides that “High-risk AI systems shall be tested for the purpose of identifying the most appropriate and targeted risk management measures”.
The Article 9 wording quoted here is the text published in OJ L series 2024/1689, apart from Article 9(1), quoted from consolidated text 02024R1689-20260727. In that consolidated text, Article 113(c), as amended by Regulation (EU) 2026/1744 of 8 July 2026, applies Chapter III, Sections 1, 2 and 3, with the exception of Article 6(5), from 2 December 2027 as regards AI systems classified as high-risk pursuant to Article 6(2) and Annex III, and from 2 August 2028 as regards AI systems classified as high-risk pursuant to Article 6(1) and Annex I. Article 9 sits in Section 2, so point (c) applies from those dates.
Consequences for a QA Function
Governance practice around these obligations is covered in AI trust and governance in quality engineering. Article 9(8) requires that “Testing shall be carried out against prior defined metrics and probabilistic thresholds that are appropriate to the intended purpose of the high-risk AI system.” The bullets below are practice built to satisfy that sentence, not additional obligations from the text. They land on timing, custody and ownership.
- Acceptance criteria - the metric and the probabilistic threshold are fixed before the run, because Article 9(8) calls them “prior defined”. Dating that definition is practice, and it makes prior easy to show.
- Evidence custody - Annex VII, point 4.3 provides that “Where relevant, and limited to what is necessary to fulfil its tasks, the notified body shall be granted full access to the training, validation, and testing data sets used”. Keeping those sets retrievable is your side of it.
- Reproducibility - Annex VII, point 4.4 says a notified body “may require that the provider supply further evidence or carry out further tests”, and that where it “is not satisfied with the tests carried out by the provider” it “shall itself directly carry out adequate tests”. Pinned model versions, pinned datasets, a recorded environment and a fixed seed are how you answer the may.
- Definition of done - for a high-risk AI system, a green run closes the ticket once the linked evidence and the reasoning behind the threshold sit alongside it.
- Mapping ownership - the line from each obligation to the specific test that evidences it has to be somebody’s job.
Ownership arrives by default, without anyone deciding it. QA holds the test IDs and the run history, so the mapping lands on QA unless someone assigns it. Scope it deliberately, well before a conformity assessment starts, where one applies to your system.
None of the following appears in the text as a requirement. These are the habits that keep an Article 9(8) record legible to a reader who was not in the room.
- Write the mapping as a table - the obligation, the tests that evidence it, the threshold, the date the threshold was set, the last run, and where the artifact lives.
- Agree a retention period with whoever owns compliance - pick a duration, apply it to logs and dataset snapshots as well as reports, and stop letting CI prune artifacts on a default window.
- Version the test data alongside the code - snapshot the validation and testing datasets used for each retained run, since Annex VII, point 4.3 names those sets by type.
- Separate exploratory runs from evidential ones - mark a small class of runs as configured, dated and retained.
- Record the real-world testing decision - Article 9(7) provides that “Testing procedures may include testing in real-world conditions in accordance with Article 60”, which is a may and not a shall. Whichever way you go, write the reasoning into the file the week the decision is made.
A testing platform covers the mechanical part of this. It runs a suite against pinned datasets, holds the environment steady between runs, and retains dated artifacts an assessor outside your team can follow.
Classification of your system and the sufficiency of the evidence you hold sit outside what any tool settles for you. Both are determinations for your own counsel and, where applicable, for the notified body assessing you. Take the legal reading from people qualified to give it.
Note: This article was written by Rahul Mishra, who works on regulatory conformity testing at TestMu AI, and produced with AI assistance under the process described in our editorial policy. Every provision cited above links to the official EU source so you can read the text rather than the summary.
Conclusion
Check your dates against the consolidated text before anything else. A plan built on 2 August 2026 for high-risk obligations is planning against a provision that no longer reads that way, and the error runs in the direction of wasted urgency rather than missed deadlines.
Then write down the metrics and thresholds while the system is still being built. Article 9(8) asks for them to exist before the testing, which is a sequencing requirement rather than a documentation one, and it is the hardest part to retrofit.
Retain evidence on the assumption that somebody outside your team will use it. An assessor who can re-run your tests needs the fixtures and the environment to have survived, not only the report.
Author
Rahul Mishra is a Lead Member of Technical Staff at TestMu AI (formerly LambdaTest), leading frontend engineering and accessibility testing across the quality engineering platform. He mentors frontend engineers, runs code reviews and sprint planning, optimizes React.js rendering performance, and makes product features accessible to users with disabilities through WCAG and ADA-compliant accessibility audits. He brings 10+ years of experience across React.js, VueJS, TypeScript, Swift, Objective-C, and AWS, with earlier work as a Technical Lead at VectoScalar Technologies. Rahul holds a B.E. in Information Technology.
Reviewer
Sirajuddin Khan is Vice President of Product Management at TestMu AI (formerly LambdaTest), where he drives the company's agentic AI product strategy, building a suite of autonomous agents that includes Agentic Browsers and Agentic Visual Testing and shifting the unit of work from test execution to autonomous outcomes. One of the company's earliest product leaders, he has owned the roadmap for the high-performance execution cloud and grew the cross-browser testing products from early adoption to market leadership. He brings over a decade of experience across SaaS, B2B, and eCommerce, with earlier product roles at Wydr and ShopClues, where his catalog and search work cut delivery SLAs and lifted seller activity. Sirajuddin holds an MBA in Information Technology from Sikkim Manipal University and a B.Tech in Computer Science Engineering from Maharshi Dayanand University.
EU AI Act Testing FAQs
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests





