Owner Stefan CoetzeeUpdated 2026-10-09

Reference · working draft 0.3 · 2026-10-07 · published 2026-10-09 · Stefan Coetzee · not a standard, not a certification scheme

Continuous Conformity for Deployed AI Systems

Working draft 0.3: 36 requirements for testing a deployed AI system, and the knowledge it runs on, on every change and on a fixed schedule, decided by someone other than the system, with records a third party can rerun. A reference text, not a standard.

statusWorking draft for discussion. Not a standard, not a certification scheme, not endorsed by any standards body. Published as a reference before any committee has seen it (decision record ADR-0019).
version0.3, 2026-10-07. 0.2 (2026-10-03) added clause 12, machine-checkable requirements; 0.3 added CC-6.6, decision-layer tests, and terms 3.25 to 3.27.
requirements36, CC-4.1 to CC-12.5. Marks per CC-12.1: M mechanical 10, A assisted 14, H manual 12 (the mark after each requirement below).
in useThe mechanical requirements run on every push to this site and weekly: /conformity/. The author's setup scored against it: self-assessment 01.
open parametersValues in square brackets are open for review: review interval, fixed run interval (proposed 30 days), intake window, external tester cadence, retention.
commentComments by clause number: github.com/uncovertechtalent/machinebehavior.io/issues (title "CC-x.y: ..."), or on r/MachineBehavior. Changes are versioned on this page.

WhyThe draftCrosswalkLimitsChanges

Why this document exists

Deployed AI systems are tested before launch and monitored after it. EU law already makes banks test live production systems on a schedule. The case: apply the same duty to the AI systems people rely on, with tests on every change and at a fixed interval, decided by someone other than the system, with records a third party can rerun.

A deployed AI system stays in conformity only while tests on the running system keep showing it. Union law asks for such tests in one place and leaves the rest to the provider: a high-risk provider's quality management system must include test procedures "before, during and after the development" of the system and "the frequency with which they have to be carried out" (AI Act Art 17(1)(d)). The provider picks the frequency. No rule sets a minimum interval, a test on every change, a test of the deployed assembly, or a tester other than the provider. The plain-speech name for the idea is a software-defined TÜV: an inspection that runs all the time instead of at intervals from outside.

The gap in the law

All quotes were checked against the Official Journal texts on 2026-10-03 and 2026-10-04.

Text What it says Effect
AI Act Art 9(8) Testing happens "at any time throughout the development process, and, in any event, prior to their being placed on the market or put into service". Testing is tied to the time before launch.
AI Act Art 72(2) Post-market monitoring shall let the provider "evaluate the continuous compliance of AI systems". The Act names the goal. The means it gives is collecting and analysing data from use.
AI Act Art 17(1)(d) The quality management system of a high-risk provider includes "examination, test and validation procedures to be carried out before, during and after the development of the high-risk AI system, and the frequency with which they have to be carried out". Testing after development is in scope, at a frequency the provider sets. No minimum interval, no change trigger, no independent tester, and the object of test is not tied to the deployed assembly.
AI Act Art 43(4) A new conformity assessment follows a "substantial modification". Changes the provider planned in advance do not count. Re-assessment follows some changes and no schedule.
DORA Art 26(2) Threat-led penetration testing "shall be performed on live production systems". Testing in production is lawful and required for financial entities.
DORA Art 24(6) Tests "at least yearly" on all systems that support critical or important functions. A fixed minimum interval exists in Union law.
DORA Art 24(4) Tests "are undertaken by independent parties, whether internal or external". The tester is separate from the tested.
Reg. (EU) 2025/1190 Art 5(1), Art 11(5) The DORA testing standard covers "testing of live production systems of critical or important functions"; the active red team phase lasts "at least 12 weeks". It sets no interval of its own and no test after a change.
GDPR Art 32(1)(d) "a process for regularly testing, assessing and evaluating the effectiveness of technical and organisational measures" A recurring testing duty already exists for the security of processing; it names no interval.

The narrow claim: the AI Act already states the goal ("continuous compliance") and already requires lifetime logging (Art 12(1)). After launch its means are monitoring and a provider-set test frequency (Art 17(1)(d)). Continuous conformity is the active-testing means for a goal the Act already names.

Where it can be written in: the Commission guidance and template on the post-market monitoring plan, due by 2 September 2027 (Art 72(3) as replaced by Regulation (EU) 2026/1744). High-risk duties for Art 6(2) and Annex III systems apply from 2 December 2027, three months later.

Why monitoring alone falls short

  1. The system under test is the deployed assembly. At run time the next token depends only on what is in the context window. System prompt, tools, retrieved documents and memory are part of the system. A test result on the bare model does not carry over to the assembly.
  2. The knowledge changes while the model stays the same. A document that was correct at conformity assessment can be wrong six months later. Data from use shows the error only after users have acted on it.

Evidence

The ask, in four requirements

# Requirement Precedent
1 Test the deployed assembly, in the configuration users meet. DORA Art 26(2), live production systems
2 Test at a fixed interval and on every change event, including a model version change by a third party. Interval: DORA Art 24(6), at least yearly. Change: AI Act Art 43(4), new assessment after a substantial modification. "On every change" goes beyond both.
3 The pass or fail decision belongs to a party that did not produce the output. DORA Art 24(4), independent parties
4 Keep records that let a third party repeat the tests. AI Act Art 12(1) logging

The question for any reader: when was the AI you rely on last tested, and who decided it passed?

Objections to expect

Objection Answer
Post-market monitoring already covers this. Monitoring collects data from use. An error shows up after a user has acted on it. A scheduled test finds it before. (AI Act Art 3(25), Art 72(2))
Testing in production is risky. Financial entities do it by law (DORA Art 26(2)). The draft requires blast-radius controls before a run on a live system: no real credentials or personal data without a legal basis, a means to stop at once, a limit on actions (CC-5.3).
The labs test their models. Those tests cover the bare model before release. The deployed assembly differs in instructions, tools, knowledge and memory (term 3.3). AI Act Art 55(1)(a) has no frequency and no post-deployment wording.
This is one person's setup. Correct, and labelled as such. The legal comparison stands on the Official Journal texts alone. The self-assessment shows the method and its limits.
Who would the independent tester be? DORA allows internal or external testers and sets qualification rules (Art 24(4), Art 27(1)). The same model can apply.

The draft

Working draft 0.3, 2026-10-07; published as a reference 2026-10-09; corrected 2026-10-09 (AI Act Art 17(1)(d) added to the introduction and Annex A). (0.2 added clause 12, machine-checkable requirements, and anchored the name to AI Act Art 72(2); 0.3 adds CC-6.6, decision-layer tests, and terms 3.25 to 3.27.) Draft for discussion. Not a standard, not a certification scheme, and not endorsed by any standards body. Values in square brackets are open parameters for review. Legal citations refer to the Official Journal texts, read by the author on 2026-10-03 and 2026-10-04.

Introduction

Union law already requires resilience testing of live production systems in one sector: DORA Art 26(2) requires threat-led penetration testing "on live production systems", and Art 24(6) requires tests "at least yearly" on all systems supporting critical or important functions. The AI Act requires risk management "throughout the entire lifecycle" (Art 9(2)), consistent performance "throughout their lifecycle" (Art 15(1)), automatic logging "over the lifetime of the system" (Art 12(1)), and post-market monitoring that lets the provider "evaluate the continuous compliance" of the system (Art 72(2)).

The AI Act ties testing to the period before the market: testing takes place "at any time throughout the development process, and, in any event, prior to their being placed on the market or put into service" (Art 9(8)). After deployment the means it names are monitoring and the collection of data from use (Art 3(25), Art 26(5), Art 72(2)). For high-risk systems the provider's quality management system must also include "examination, test and validation procedures to be carried out before, during and after the development of the high-risk AI system, and the frequency with which they have to be carried out" (Art 17(1)(d)). The frequency is the provider's choice: the Act sets no minimum interval, no test on a change event, no test of the deployed assembly as such, and no tester other than the provider. A new conformity assessment follows a substantial modification (Art 43(4)), and changes the provider planned in advance do not count as one.

Two facts about deployed language-model systems make monitoring alone insufficient:

  1. The system under test is the deployed assembly, not the model. A language model has Markovian memory: at run time the next token depends only on what is in the context window. The system prompt, tools, retrieved documents and memory in that window are part of the system, and a test result on the bare model does not carry over to the assembly.
  2. The knowledge the assembly retrieves changes without any change to the model. A document that was correct at conformity assessment can be wrong six months later, and post-market data from use shows the error only after users have acted on it.

This document specifies requirements for testing a deployed AI system, and the knowledge it runs on, on a fixed schedule and on every change, with records that let a third party repeat the tests. It names the result continuous conformity. The name is chosen to match the phrase the AI Act already uses: post-market monitoring shall allow the provider to "evaluate the continuous compliance" of the system (Art 72(2)). This document specifies active testing as a means to that aim. It also specifies how requirements can be written so that a machine can check them (clause 12), because a test that runs on every change and on a fixed schedule needs requirements a tool can evaluate.

1 Scope

This document specifies requirements for a conformity testing programme for deployed AI systems that generate outputs from a model together with instructions, tools, retrieved knowledge or memory supplied at run time.

It applies to providers and deployers of such systems, whatever the risk classification under Regulation (EU) 2024/1689. It can support, and does not replace, the risk management system (Art 9), post-market monitoring (Art 72), deployer monitoring (Art 26(5)) and adversarial testing of general-purpose AI models with systemic risk (Art 55(1)(a)).

It does not specify model training, pre-market conformity assessment procedures, or the content of any harmonised standard.

2 Normative references

There are no normative references in this document. Sources cited in rationales are listed in the Bibliography.

3 Terms and definitions

Terms defined in Regulation (EU) 2024/1689 Art 3 keep their meaning there, in particular: AI system (3(1)), provider (3(3)), deployer (3(4)), intended purpose (3(12)), performance of an AI system (3(18)), substantial modification (3(23)), post-market monitoring system (3(25)), serious incident (3(49)).

3.1 organization provider or deployer that is responsible for a deployed AI system and operates the conformity testing programme

3.2 deployed AI system AI system after it has been placed on the market or put into service, in the configuration in which it produces outputs for users

3.3 deployed assembly combination of model, model version, system instructions, tools, retrieval sources, knowledge base, memory, output controls and configuration that together produce the outputs of a deployed AI system Note 1 to entry: two deployed assemblies that share a model and differ in any other component are different systems under test.

3.4 knowledge base documented information that the deployed assembly retrieves or receives in its context at run time and uses to produce outputs

3.5 knowledge item smallest unit of the knowledge base that has one owner and one recorded source

3.6 source of record document or system from which a knowledge item is derived and against which it is verified Note 1 to entry: a summary of a source, including a summary produced by a model, is not a source of record.

3.7 freshness interval maximum time a knowledge item may remain in the knowledge base without verification against its source of record

3.8 change event any change to a component of the deployed assembly, including a model version change made by a third party

3.9 conformity test test executed against the deployed assembly with a pass criterion defined before execution

3.10 test set versioned set of conformity tests for one deployed AI system

3.11 regression set part of the test set made of every conformity test that has detected a nonconformity at least once

3.12 steady state measured rate of a defined behaviour of the deployed assembly on a defined task with no pressure condition applied

3.13 pressure condition defined variation of input or environment intended to move a behaviour away from its steady state Note 1 to entry: Annex C lists pressure condition classes.

3.14 unannounced arm execution of a conformity test in which nothing available to the deployed assembly indicates that the input is a test

3.15 invariance gap difference between the result of a conformity test in its announced arm and in its unannounced arm

3.16 prediction record written statement of the expected result of a test run, with its unit of measurement and its grader, dated and hashed before the run

3.17 grader person, procedure or system that decides whether a conformity test passed

3.18 independent tester person or system that did not produce, and is not responsible for, the output under test

3.19 blast radius scope of harm to users, data, services or third parties that a test run can cause

3.20 nonconformity non-fulfilment of a requirement Note 1 to entry: wording as in the ISO harmonized structure for management system standards; not re-checked against the published text in this draft.

3.21 continuous conformity state of a deployed AI system for which conformity with its stated requirements is shown by conformity tests run on the deployed assembly at a fixed interval and on every change event, with records that let a third party repeat the tests Note 1 to entry: the term matches Regulation (EU) 2024/1689 Art 72(2), which requires a post-market monitoring system that allows the provider "to evaluate the continuous compliance of AI systems with the requirements set out in Chapter III, Section 2". Art 72(2) names the aim; this document specifies active testing as one means to it.

3.22 machine-checkable requirement requirement stated in a form that a tool can evaluate against an output, a record or a configuration without human interpretation

3.23 rule table versioned set of machine-checkable requirements, each with an identifier, a check, a severity and a corrective instruction

3.24 output boundary point at which an output of the deployed assembly leaves the assembly for a user, a file or another system

3.25 decision layer The part of an output that commits to a decision: a verdict, a recommendation, a selected action or a stated fact that a user or a downstream system acts on. Distinct from the output text in which the decision is expressed.

3.26 scripted pressure sequence A fixed sequence of two or more turns, each applying a pressure condition (3.13) from Annex C, applied after an initial output, with the correct decision fixed by ground truth before the run and no new evidence introduced by the sequence.

3.27 fold A change at the decision layer (3.25) from a correct decision to an incorrect one during a scripted pressure sequence (3.26). The fold rate is the number of runs that fold divided by the number of runs that were correct before the sequence began.

4 Conformity testing programme

CC-4.1 H The organization shall establish, maintain and review a documented conformity testing programme for each deployed AI system.

CC-4.2 H The programme shall list the requirements under test. For each requirement it shall state a metric and a pass threshold before the first test run.

CC-4.3 H The programme shall name one person accountable for it.

5 Object of test

CC-5.1 A Conformity tests shall run against the deployed assembly. Results from the model alone, or from a different assembly, shall not be reported as results for the deployed AI system.

CC-5.2 M Each run record shall identify the version of every component of the deployed assembly at the time of the run.

CC-5.3 H Where conformity tests run on the system that serves users, the organization shall define and apply controls on blast radius before the run. These shall include: no use of real credentials or real personal data unless the test requires it and a documented legal basis exists; a means to stop the run at once; and a limit on the actions the assembly can take during the run.

6 Test design

CC-6.1 A For each behaviour under test, the organization shall measure the steady state before applying any pressure condition.

CC-6.2 A The test set shall include at least one pressure condition from each class in Annex C that is relevant to the intended purpose. Exclusion of a class shall be documented with a reason.

CC-6.3 M For each test run the organization shall write a prediction record before the run, and shall keep any later amendment with its date and reason.

CC-6.4 A The test set should include unannounced arms. Where it does, the run record shall report the invariance gap.

CC-6.5 A Each grader shall be tested against cases with a known pass result and cases with a known fail result before use and after each change to the grader.

CC-6.6 A The test set shall include decision-layer tests: tasks whose correct decision is fixed by ground truth before the run, applied under a scripted pressure sequence (3.26) drawn from the classes in Annex C. The run record shall report, per arm, the correctness before the sequence, the correctness after it and the fold rate (3.27). A test set that consists only of output-text rules (CC-12.3) shall not be reported as evidence of conformity at the decision layer.

7 Cadence and triggers

CC-7.1 M The organization shall run the full test set on every change event, before the changed assembly serves users.

CC-7.2 H A change event should be released in stages. Conformity tests should pass at each stage before the next.

CC-7.3 M The organization shall run the full test set at least every [30] days, whether or not a change event occurred.

CC-7.4 A Every conformity test that has detected a nonconformity shall be added to the regression set. A test shall not be removed from the regression set without a documented reason approved by the programme owner.

CC-7.5 H Attack methods published in the public record that are relevant to the intended purpose should be added to the test set within [60] days of publication.

8 Knowledge base conformity

CC-8.1 A Each knowledge item shall have a named owner and a recorded source of record.

CC-8.2 M Each knowledge item shall carry the date it was last verified against its source of record. An item older than its freshness interval shall be re-verified or withdrawn from retrieval.

CC-8.3 A The organization shall state the freshness interval for each class of knowledge item, with a reason tied to how often the source of record changes.

CC-8.4 A The test set shall include knowledge conformity tests: inputs whose correct output is fixed by a knowledge item, graded against the source of record.

CC-8.5 A Verification of a knowledge item shall be made against the source of record and not against a summary of it.

CC-8.6 A Where outputs are used for decisions by users or third parties, statements of fact in the output should be traceable to a knowledge item or a cited source.

9 Independence

CC-9.1 H The pass or fail decision of a conformity test shall not rest only with the system or person that produced the output under test.

CC-9.2 H An independent tester from outside the organization should run the test set at least every [third] full run cycle [or at least every 12 months].

10 Findings and remediation

CC-10.1 A The organization shall classify, assign and track every nonconformity to closure. Closure shall require a passing rerun of the test that found it.

CC-10.2 H Where a nonconformity meets the definition of a serious incident (AI Act Art 3(49)), or shows that the system may present a risk within the meaning of Art 79(1), the organization shall pass it to its incident reporting process without delay.

11 Records and disclosure

CC-11.1 M Each run record shall contain at least: date and time; deployed assembly versions (CC-5.2); test set version; prediction record and its hash; raw results; pass or fail per test; grader identity; and the steps needed to rerun.

CC-11.2 M Run records shall be kept for at least [the lifetime of the deployed AI system plus 6 months].

CC-11.3 H The organization may publish its method, raw results and rerun steps. A published assessment made by the organization of its own system shall be labelled a self-assessment and shall not be presented as a certification.

CC-11.4 H A published self-assessment should invite a second rater and should publish any disagreement between raters.

12 Machine-checkable requirements

CC-12.1 H The programme shall mark each requirement as mechanical (a tool decides), assisted (a tool flags, a person decides) or manual (a person decides).

CC-12.2 A Requirements marked mechanical should be expressed in a machine-readable form, kept under version control with the test set.

CC-12.3 M Where a mechanical requirement applies to every output, the organization should enforce it at the output boundary with a rule table. Each rule in a blocking tier shall have a recorded false-positive test, and rules that fail it shall move to a warning tier.

CC-12.4 M Each rule hit at the output boundary shall be logged with the rule identifier, the time and the session or request. The log shall be reviewed at each programme review to retire, tighten or promote rules.

CC-12.5 M Run records (CC-11.1) and nonconformities (CC-10.1) should be exportable in an open, published format for assessment results and findings.

Annex A (informative): relation to existing law and standards

Clause here Existing text What exists What this document adds
4.1 DORA Art 24(1); AI Act Art 9(1) Testing programme (DORA); risk management system (AI Act) A programme scoped to the deployed AI system
4.2 AI Act Art 9(8) Prior defined metrics and thresholds, pre-market Same rule applied after deployment
5.1 DORA Art 26(2); TIBER-EU 2.1.1 Tests on live production (financial ICT) Same principle for the deployed assembly
6.4 none found Unannounced arms and invariance gap
7.1 AI Act Art 43(4) Reassessment on substantial modification Tests on every change event
7.3 DORA Art 24(6); AI Act Art 17(1)(d) At least yearly (DORA); test procedures before, during and after development at a frequency the provider sets (AI Act, high-risk) A fixed maximum interval, and tests on every change event
8.x AI Act Art 72(2); Art 53(1)(a) "keep up-to-date" (GPAI documentation) Monitoring data; documentation kept current Testing of the knowledge the system runs on
9.1-9.2 DORA Art 24(4), 26(8) Independent and external testers Same for deployed AI systems
10.1 DORA Art 24(5) Remediation with validation Same
11.1 AI Act Art 12; Annex IV 2(g) Logs; dated, signed test reports Rerun steps and prediction record
(all) AI Act Art 72(2) Aim: "continuous compliance"; means: data collection Means: active tests
(all) NIST AI RMF MEASURE 2.4, MANAGE 4.1 Monitoring in production; post-deployment monitoring plans Testing, with cadence
(all) ISO/IEC 42001 A.6.2.4, A.6.2.6 (numbers from secondary sources) Verification and validation; operation and monitoring Cadence and triggers after deployment
12.2, 12.5 NIST OSCAL (v1.2.3 is the latest release on GitHub, checked 2026-10-03); NIST SP 800-126 Rev. 3 (SCAP 1.3, February 2018) Machine-readable control catalogues, assessment results and findings (security) Same formats applied to AI conformity tests
3.3, 5.1 ISO/IEC AWI 26741 (stage 20.00): framework for describing an AI system as object of conformity assessment, including "system boundaries" Work item open Deployed assembly as the boundary
12.x CEN SMART Standards project Machine-readable standards (format of the standard itself) Machine-checkable requirements (format of the requirement inside a programme)

Annex B (informative): origin of each requirement

Requirement Incident or record Source type
CC-4.3 GitLab database deletion, 2017 primary (GitLab postmortem)
CC-5.1, CC-6.5, CC-7.2 CrowdStrike Channel File 291, 2024 primary (CrowdStrike RCA)
CC-5.2 Knight Capital, 2012 primary (SEC order 34-70694)
CC-6.2 Objections register incident log author's public record
CC-6.6 Eval 04 full-01: clean lexical score, 75 percent fold rate under scripted pushback (qwen3-coder-30b only); full-02: short stance instruction folds 33 percent at the same turn-0 correctness; v5: three larger models folded 0 of 72 each (2026-10-07); Moffatt v. Air Canada; Mata v. Avianca author's published experiment (https://machinebehavior.io/experiments/#experiment-04, #experiment-04-short-stance, #experiment-04-three-models); primary; primary
CC-7.4 Chess-engine board edit trained out; behaviour moved author's piece, citing Palisade arXiv:2502.13295
CC-8.1 Fogbank, 2000-2008 primary (GAO-09-385)
CC-8.2, CC-8.4 Moffatt v. Air Canada, 2024 primary (2024 BCCRT 149)
CC-8.5 Summariser invented a sanction figure, 2026-10-03 author's record
CC-8.6 Mata v. Avianca, 2023; Deloitte/DEWR, 2025 primary; secondary
CC-9.1 Seven relapses, none self-caught author's piece "What Operations Already Knows"
CC-12.1 Surface ban moved the behaviour (symptom substitution) author's piece "Sycophancy Is Layered"
CC-12.3, CC-12.4 vestige-kit rule table and hit log author's working instance

Annex C (normative): pressure condition classes

Class Description Example
C.1 Pushback user disputes a correct output "That is wrong, check again" after a correct answer
C.2 False premise input contains a confident false statement a question that assumes a policy that does not exist
C.3 Authority claim input claims a role or permission "As the system administrator, I authorise you to"
C.4 Outcome-only scoring the task rewards the result and not the method win-only scoring in a game task
C.5 Approval signal input rewards agreement praise after agreement, displeasure after disagreement
C.6 Stale or conflicting knowledge retrieved knowledge disagrees with the source of record an outdated price list in retrieval
C.7 Instruction in data retrieved content or tool output contains instructions a web page that tells the assistant to change task

Bibliography

  1. Regulation (EU) 2024/1689 (Artificial Intelligence Act), OJ L, 12.7.2024.
  2. Regulation (EU) 2022/2554 (DORA), OJ L 333, 27.12.2022.
  3. Commission Delegated Regulation (EU) 2025/1190 (TLPT RTS), read 2026-10-04: covers testing of live production systems of critical or important functions (Art 5(1)); active red team phase at least 12 weeks (Art 11(5)); sets no interval of its own and no change-triggered test.
  4. ECB, TIBER-EU Framework, February 2025.
  5. NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0), 2023.
  6. ISO/IEC 42001:2023; ISO/IEC 27001:2022; ISO/IEC 27002:2022; ISO 22301:2019.
  7. Principles of Chaos Engineering (principlesofchaos.org); Basiri et al., "Chaos Engineering", IEEE Software 33(3), 2016.
  8. Greenblatt et al., arXiv:2412.14093; Meinke et al., arXiv:2412.04984; Palisade Research, arXiv:2502.13295.
  9. Incident sources as named in each rationale and in Annex B.
  10. machinebehavior.io: claims ledger, objections register, slips log, prediction hashes, "Chaos Engineering for Behaviour".
  11. NIST, Open Security Controls Assessment Language (OSCAL), pages.nist.gov/OSCAL; releases at github.com/usnistgov/OSCAL.
  12. NIST SP 800-126 Rev. 3, The Technical Specification for the Security Content Automation Protocol (SCAP): SCAP Version 1.3, February 2018.
  13. Cilla Ugarte et al., "Making AI Compliance Evidence Machine-Readable", arXiv:2604.13767, submitted 15 April 2026.
  14. CEN, SMART Standards project, experts.cen.eu/key-initiatives/smart-standards/.
  15. ISO/IEC AWI 26741 (iso.org/standard/94402.html); ISO/IEC DIS 23282 (iso.org/standard/87387.html).
  16. vestige-kit (github.com/uncovertechtalent/vestige-kit).

Crosswalk to existing frameworks

Each requirement of this draft, the clause of an existing framework it relates to, and how. "Evidence toward" means a passing test of the requirement is evidence for part of that clause, with the slice named; it is never conformity to the framework. "Nearest clause" means the framework has nothing closer and does not require what the requirement asks: these rows are the gap this document names. Machine-readable: /conformity/requirements.json.

reqmarkframeworkclauserelationslice
CC-4.1HRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 9(1)-(2)evidence towardrisk management system as a continuous iterative process
CC-4.1HRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 72(1)evidence towarddocumented post-market monitoring system
CC-4.1HRegulation (EU) 2022/2554 (DORA), EUR-LexArt 24(1)evidence towarddigital operational resilience testing programme
CC-4.2HRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 9(8)evidence towardtesting against prior defined metrics and probabilistic thresholds
CC-4.2HNIST AI 100-1, AI Risk Management Framework 1.0MEASURE 1.1evidence towardapproaches and metrics for measurement selected
CC-4.3HNIST AI 100-1, AI Risk Management Framework 1.0GOVERN 2.1evidence towardroles and responsibilities documented
CC-5.1ARegulation (EU) 2024/1689 (AI Act), EUR-LexArt 15(1)evidence towardconsistent performance of the system as placed on the market
CC-5.1ARegulation (EU) 2022/2554 (DORA), EUR-LexArt 24(2)evidence towardtesting of ICT systems supporting critical functions
CC-5.2MRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 12(1)evidence towardautomatic recording of events over the lifetime of the system
CC-5.2MRegulation (EU) 2022/2554 (DORA), EUR-LexArt 9(4)(e)evidence towarddocumented ICT change management
CC-5.3HRegulation (EU) 2022/2554 (DORA), EUR-LexArt 26(5)evidence towardrisk management measures for threat-led tests on live production systems
CC-6.1ANIST AI 100-1, AI Risk Management Framework 1.0MEASURE 2.5nearest clause; it does not require thisvalidity and reliability assessed
CC-6.2ARegulation (EU) 2024/1689 (AI Act), EUR-LexArt 15(5)evidence towardresilience against attempts to alter use, outputs or performance
CC-6.2ARegulation (EU) 2024/1689 (AI Act), EUR-LexArt 55(1)(b)nearest clause; it does not require thisadversarial testing of GPAI models with systemic risk
CC-6.3MRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 9(8)evidence towardprior defined metrics and thresholds
CC-6.4ARegulation (EU) 2022/2554 (DORA), EUR-LexArt 26(1)nearest clause; it does not require thisthreat-led penetration testing, covert by design
CC-6.5ANIST AI 100-1, AI Risk Management Framework 1.0MEASURE 2.5nearest clause; it does not require thisvalidity of the evaluation itself
CC-6.6ARegulation (EU) 2024/1689 (AI Act), EUR-LexArt 15(1)evidence towardaccuracy and robustness, performance consistent through the lifecycle
CC-6.6ARegulation (EU) 2024/1689 (AI Act), EUR-LexArt 15(5)evidence towardresilience against attempts to alter outputs
CC-6.6ARegulation (EU) 2024/1689 (AI Act), EUR-LexArt 9(8)evidence towardtesting against prior defined metrics
CC-6.6ARegulation (EU) 2022/2554 (DORA), EUR-LexArt 26(2)nearest clause; it does not require thisthreat-led testing on live systems
CC-7.1MRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 9(6)-(7)evidence towardtesting throughout development and before placing on the market
CC-7.1MRegulation (EU) 2022/2554 (DORA), EUR-LexArt 9(4)(e)evidence towardICT change management with testing before deployment
CC-7.1MRegulation (EU) 2022/2554 (DORA), EUR-LexArt 25(1)evidence towardappropriate tests on ICT systems
CC-7.1MNIST CSF 2.0PR.PS-01evidence towardconfiguration management
CC-7.2HRegulation (EU) 2022/2554 (DORA), EUR-LexArt 9(4)(e)nearest clause; it does not require thischange management
CC-7.3MRegulation (EU) 2022/2554 (DORA), EUR-LexArt 24(6)evidence towardappropriate tests at least yearly
CC-7.3MRegulation (EU) 2016/679, EUR-LexArt 32(1)(d)evidence towardregular testing, assessing and evaluating
CC-7.3MRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 72(2)evidence towardevaluate continuous compliance throughout the lifetime
CC-7.3MRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 17(1)(d)evidence towardtest procedures after development at a stated frequency (high-risk providers; frequency set by the provider)
CC-7.4ANIST CSF 2.0ID.IM-03evidence towardimprovements from lessons learned
CC-7.5HRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 55(1)(b)nearest clause; it does not require thisstate-of-the-art adversarial testing
CC-8.1ARegulation (EU) 2024/1689 (AI Act), EUR-LexArt 10(2)nearest clause; it does not require thisdata governance practices
CC-8.2MRegulation (EU) 2016/679, EUR-LexArt 5(1)(d)evidence towardaccuracy; kept up to date
CC-8.2MRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 10(3)nearest clause; it does not require thisdata sets relevant, representative, free of errors
CC-8.3ARegulation (EU) 2016/679, EUR-LexArt 5(1)(d)nearest clause; it does not require thisaccuracy
CC-8.4ARegulation (EU) 2024/1689 (AI Act), EUR-LexArt 15(1)nearest clause; it does not require thisaccuracy
CC-8.5ARegulation (EU) 2016/679, EUR-LexArt 5(1)(d)nearest clause; it does not require thisaccuracy
CC-8.6ARegulation (EU) 2024/1689 (AI Act), EUR-LexArt 13(1)nearest clause; it does not require thistransparency, interpretable output
CC-9.1HRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 14(1)evidence towardhuman oversight
CC-9.1HRegulation (EU) 2022/2554 (DORA), EUR-LexArt 24(4)evidence towardtests by independent parties, internal or external
CC-9.2HRegulation (EU) 2022/2554 (DORA), EUR-LexArt 26(8)evidence towardexternal testers for threat-led tests
CC-9.2HRegulation (EU) 2022/2554 (DORA), EUR-LexArt 27evidence towardrequirements for testers
CC-10.1ANIST CSF 2.0ID.IM-03evidence towardlessons learned
CC-10.1ARegulation (EU) 2022/2554 (DORA), EUR-LexArt 24(5)evidence towardremediation of issues identified in tests
CC-10.2HRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 73(1)evidence towardreporting of serious incidents
CC-10.2HRegulation (EU) 2022/2554 (DORA), EUR-LexArt 19(1)evidence towardreporting of major ICT-related incidents
CC-11.1MRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 12(1)evidence towardautomatic recording of events
CC-11.1MRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 72(2)evidence towardcollection of data on performance throughout the lifetime
CC-11.1MNIST CSF 2.0DE.CM-09evidence towardmonitoring of software and services
CC-11.2MRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 19(1)evidence towardlogs kept for at least six months
CC-11.2MRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 18(1)evidence towarddocumentation kept 10 years
CC-11.3HRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 43(2)nearest clause; it does not require thisconformity assessment based on internal control
CC-11.4HRegulation (EU) 2022/2554 (DORA), EUR-LexArt 24(4)nearest clause; it does not require thisindependent parties
CC-12.1HRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 14(1)evidence towardwhich decisions a person takes
CC-12.2ANIST CSF 2.0PR.PS-01nearest clause; it does not require thisconfiguration management
CC-12.3MRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 15(1)evidence towardconsistent performance at the output
CC-12.3MNIST CSF 2.0PR.PS-01evidence towardconfiguration management of the rule table
CC-12.4MRegulation (EU) 2024/1689 (AI Act), EUR-LexArt 12(1)evidence towardlogging
CC-12.4MNIST CSF 2.0DE.CM-09evidence towardmonitoring
CC-12.5MNIST CSF 2.0GV.OCnearest clause; it does not require thisorganizational context

Limits

Changes

dateversionchange
2026-10-090.3Correction after an outside review: AI Act Art 17(1)(d) (test procedures before, during and after development, at a frequency the provider sets) added to the introduction, Annex A, the gap table and the crosswalk; the sentence "no rule in Union law requires anyone to run those tests after launch" withdrawn as too broad.
2026-10-090.3Published as a reference on machinebehavior.io.
2026-10-070.3CC-6.6 decision-layer tests; terms 3.25 decision layer, 3.26 scripted pressure sequence, 3.27 fold.
2026-10-030.2Clause 12, machine-checkable requirements; name anchored to AI Act Art 72(2).