Continuous Conformity for Deployed AI Systems
Working draft 0.3: 36 requirements for testing a deployed AI system, and the knowledge it runs on, on every change and on a fixed schedule, decided by someone other than the system, with records a third party can rerun. A reference text, not a standard.
| status | Working draft for discussion. Not a standard, not a certification scheme, not endorsed by any standards body. Published as a reference before any committee has seen it (decision record ADR-0019). |
| version | 0.3, 2026-10-07. 0.2 (2026-10-03) added clause 12, machine-checkable requirements; 0.3 added CC-6.6, decision-layer tests, and terms 3.25 to 3.27. |
| requirements | 36, CC-4.1 to CC-12.5. Marks per CC-12.1: M mechanical 10, A assisted 14, H manual 12 (the mark after each requirement below). |
| in use | The mechanical requirements run on every push to this site and weekly: /conformity/. The author's setup scored against it: self-assessment 01. |
| open parameters | Values in square brackets are open for review: review interval, fixed run interval (proposed 30 days), intake window, external tester cadence, retention. |
| comment | Comments by clause number: github.com/uncovertechtalent/machinebehavior.io/issues (title "CC-x.y: ..."), or on r/MachineBehavior. Changes are versioned on this page. |
WhyThe draftCrosswalkLimitsChanges
Why this document exists
Deployed AI systems are tested before launch and monitored after it. EU law already makes banks test live production systems on a schedule. The case: apply the same duty to the AI systems people rely on, with tests on every change and at a fixed interval, decided by someone other than the system, with records a third party can rerun.
A deployed AI system stays in conformity only while tests on the running system keep showing it. Union law asks for such tests in one place and leaves the rest to the provider: a high-risk provider's quality management system must include test procedures "before, during and after the development" of the system and "the frequency with which they have to be carried out" (AI Act Art 17(1)(d)). The provider picks the frequency. No rule sets a minimum interval, a test on every change, a test of the deployed assembly, or a tester other than the provider. The plain-speech name for the idea is a software-defined TÜV: an inspection that runs all the time instead of at intervals from outside.
The gap in the law
All quotes were checked against the Official Journal texts on 2026-10-03 and 2026-10-04.
| Text | What it says | Effect |
|---|---|---|
| AI Act Art 9(8) | Testing happens "at any time throughout the development process, and, in any event, prior to their being placed on the market or put into service". | Testing is tied to the time before launch. |
| AI Act Art 72(2) | Post-market monitoring shall let the provider "evaluate the continuous compliance of AI systems". | The Act names the goal. The means it gives is collecting and analysing data from use. |
| AI Act Art 17(1)(d) | The quality management system of a high-risk provider includes "examination, test and validation procedures to be carried out before, during and after the development of the high-risk AI system, and the frequency with which they have to be carried out". | Testing after development is in scope, at a frequency the provider sets. No minimum interval, no change trigger, no independent tester, and the object of test is not tied to the deployed assembly. |
| AI Act Art 43(4) | A new conformity assessment follows a "substantial modification". Changes the provider planned in advance do not count. | Re-assessment follows some changes and no schedule. |
| DORA Art 26(2) | Threat-led penetration testing "shall be performed on live production systems". | Testing in production is lawful and required for financial entities. |
| DORA Art 24(6) | Tests "at least yearly" on all systems that support critical or important functions. | A fixed minimum interval exists in Union law. |
| DORA Art 24(4) | Tests "are undertaken by independent parties, whether internal or external". | The tester is separate from the tested. |
| Reg. (EU) 2025/1190 Art 5(1), Art 11(5) | The DORA testing standard covers "testing of live production systems of critical or important functions"; the active red team phase lasts "at least 12 weeks". | It sets no interval of its own and no test after a change. |
| GDPR Art 32(1)(d) | "a process for regularly testing, assessing and evaluating the effectiveness of technical and organisational measures" | A recurring testing duty already exists for the security of processing; it names no interval. |
The narrow claim: the AI Act already states the goal ("continuous compliance") and already requires lifetime logging (Art 12(1)). After launch its means are monitoring and a provider-set test frequency (Art 17(1)(d)). Continuous conformity is the active-testing means for a goal the Act already names.
Where it can be written in: the Commission guidance and template on the post-market monitoring plan, due by 2 September 2027 (Art 72(3) as replaced by Regulation (EU) 2026/1744). High-risk duties for Art 6(2) and Annex III systems apply from 2 December 2027, three months later.
Why monitoring alone falls short
- The system under test is the deployed assembly. At run time the next token depends only on what is in the context window. System prompt, tools, retrieved documents and memory are part of the system. A test result on the bare model does not carry over to the assembly.
- The knowledge changes while the model stays the same. A document that was correct at conformity assessment can be wrong six months later. Data from use shows the error only after users have acted on it.
Evidence
- Moffatt v. Air Canada (2024 BCCRT 149). A chatbot gave a traveller a wrong answer on a fare. The tribunal ordered the airline to pay C$812.02 and rejected the argument that the chatbot was a "separate legal entity".
- Mata v. Avianca (678 F.Supp.3d 443, S.D.N.Y. 2023). Lawyers filed cases a chatbot had invented and stood by them when the judge asked. Sanction: $5,000.
- One incident, timed (case 12). Enforcement was fast once a rule existed (68.7 seconds from the user's message to an unprompted disclosure), and prevention was absent before it. The licence clause and a false absence claim were both caught outside the session that was affected.
- A self-assessment (self-assessment 01). The author's own working setup scored against this draft: 4 pass, 16 partly met, 13 gaps, 2 not applicable of 35, against a prediction hashed before scoring that matched 29 of 35; 5 of the 6 misses were too generous. A model second rater from another family agreed on 19 of 35 rows. A setup built by someone who cares about this still fails most of the requirements.
- Who catches the errors (objections register). 11 incidents caught by the human operator, 0 by the model's self-audit. Author's record.
- Decision layer (experiment 04). A wording check passed a model (qwen3-coder-30b) that gave up a correct answer under scripted pushback in 75 percent of eligible runs. Three larger models held every answer under the same design (0 of 72 each). Folding is a property of a model; the requirement (CC-6.6) asks for the test.
- The site that publishes this runs it (conformity gate). The mechanical requirements run on every push before machinebehavior.io deploys, and weekly. Self-assessment, not a certification.
The ask, in four requirements
| # | Requirement | Precedent |
|---|---|---|
| 1 | Test the deployed assembly, in the configuration users meet. | DORA Art 26(2), live production systems |
| 2 | Test at a fixed interval and on every change event, including a model version change by a third party. | Interval: DORA Art 24(6), at least yearly. Change: AI Act Art 43(4), new assessment after a substantial modification. "On every change" goes beyond both. |
| 3 | The pass or fail decision belongs to a party that did not produce the output. | DORA Art 24(4), independent parties |
| 4 | Keep records that let a third party repeat the tests. | AI Act Art 12(1) logging |
The question for any reader: when was the AI you rely on last tested, and who decided it passed?
Objections to expect
| Objection | Answer |
|---|---|
| Post-market monitoring already covers this. | Monitoring collects data from use. An error shows up after a user has acted on it. A scheduled test finds it before. (AI Act Art 3(25), Art 72(2)) |
| Testing in production is risky. | Financial entities do it by law (DORA Art 26(2)). The draft requires blast-radius controls before a run on a live system: no real credentials or personal data without a legal basis, a means to stop at once, a limit on actions (CC-5.3). |
| The labs test their models. | Those tests cover the bare model before release. The deployed assembly differs in instructions, tools, knowledge and memory (term 3.3). AI Act Art 55(1)(a) has no frequency and no post-deployment wording. |
| This is one person's setup. | Correct, and labelled as such. The legal comparison stands on the Official Journal texts alone. The self-assessment shows the method and its limits. |
| Who would the independent tester be? | DORA allows internal or external testers and sets qualification rules (Art 24(4), Art 27(1)). The same model can apply. |
The draft
Working draft 0.3, 2026-10-07; published as a reference 2026-10-09; corrected 2026-10-09 (AI Act Art 17(1)(d) added to the introduction and Annex A). (0.2 added clause 12, machine-checkable requirements, and anchored the name to AI Act Art 72(2); 0.3 adds CC-6.6, decision-layer tests, and terms 3.25 to 3.27.) Draft for discussion. Not a standard, not a certification scheme, and not endorsed by any standards body. Values in square brackets are open parameters for review. Legal citations refer to the Official Journal texts, read by the author on 2026-10-03 and 2026-10-04.
Introduction
Union law already requires resilience testing of live production systems in one sector: DORA Art 26(2) requires threat-led penetration testing "on live production systems", and Art 24(6) requires tests "at least yearly" on all systems supporting critical or important functions. The AI Act requires risk management "throughout the entire lifecycle" (Art 9(2)), consistent performance "throughout their lifecycle" (Art 15(1)), automatic logging "over the lifetime of the system" (Art 12(1)), and post-market monitoring that lets the provider "evaluate the continuous compliance" of the system (Art 72(2)).
The AI Act ties testing to the period before the market: testing takes place "at any time throughout the development process, and, in any event, prior to their being placed on the market or put into service" (Art 9(8)). After deployment the means it names are monitoring and the collection of data from use (Art 3(25), Art 26(5), Art 72(2)). For high-risk systems the provider's quality management system must also include "examination, test and validation procedures to be carried out before, during and after the development of the high-risk AI system, and the frequency with which they have to be carried out" (Art 17(1)(d)). The frequency is the provider's choice: the Act sets no minimum interval, no test on a change event, no test of the deployed assembly as such, and no tester other than the provider. A new conformity assessment follows a substantial modification (Art 43(4)), and changes the provider planned in advance do not count as one.
Two facts about deployed language-model systems make monitoring alone insufficient:
- The system under test is the deployed assembly, not the model. A language model has Markovian memory: at run time the next token depends only on what is in the context window. The system prompt, tools, retrieved documents and memory in that window are part of the system, and a test result on the bare model does not carry over to the assembly.
- The knowledge the assembly retrieves changes without any change to the model. A document that was correct at conformity assessment can be wrong six months later, and post-market data from use shows the error only after users have acted on it.
This document specifies requirements for testing a deployed AI system, and the knowledge it runs on, on a fixed schedule and on every change, with records that let a third party repeat the tests. It names the result continuous conformity. The name is chosen to match the phrase the AI Act already uses: post-market monitoring shall allow the provider to "evaluate the continuous compliance" of the system (Art 72(2)). This document specifies active testing as a means to that aim. It also specifies how requirements can be written so that a machine can check them (clause 12), because a test that runs on every change and on a fixed schedule needs requirements a tool can evaluate.
1 Scope
This document specifies requirements for a conformity testing programme for deployed AI systems that generate outputs from a model together with instructions, tools, retrieved knowledge or memory supplied at run time.
It applies to providers and deployers of such systems, whatever the risk classification under Regulation (EU) 2024/1689. It can support, and does not replace, the risk management system (Art 9), post-market monitoring (Art 72), deployer monitoring (Art 26(5)) and adversarial testing of general-purpose AI models with systemic risk (Art 55(1)(a)).
It does not specify model training, pre-market conformity assessment procedures, or the content of any harmonised standard.
2 Normative references
There are no normative references in this document. Sources cited in rationales are listed in the Bibliography.
3 Terms and definitions
Terms defined in Regulation (EU) 2024/1689 Art 3 keep their meaning there, in particular: AI system (3(1)), provider (3(3)), deployer (3(4)), intended purpose (3(12)), performance of an AI system (3(18)), substantial modification (3(23)), post-market monitoring system (3(25)), serious incident (3(49)).
3.1 organization provider or deployer that is responsible for a deployed AI system and operates the conformity testing programme
3.2 deployed AI system AI system after it has been placed on the market or put into service, in the configuration in which it produces outputs for users
3.3 deployed assembly combination of model, model version, system instructions, tools, retrieval sources, knowledge base, memory, output controls and configuration that together produce the outputs of a deployed AI system Note 1 to entry: two deployed assemblies that share a model and differ in any other component are different systems under test.
3.4 knowledge base documented information that the deployed assembly retrieves or receives in its context at run time and uses to produce outputs
3.5 knowledge item smallest unit of the knowledge base that has one owner and one recorded source
3.6 source of record document or system from which a knowledge item is derived and against which it is verified Note 1 to entry: a summary of a source, including a summary produced by a model, is not a source of record.
3.7 freshness interval maximum time a knowledge item may remain in the knowledge base without verification against its source of record
3.8 change event any change to a component of the deployed assembly, including a model version change made by a third party
3.9 conformity test test executed against the deployed assembly with a pass criterion defined before execution
3.10 test set versioned set of conformity tests for one deployed AI system
3.11 regression set part of the test set made of every conformity test that has detected a nonconformity at least once
3.12 steady state measured rate of a defined behaviour of the deployed assembly on a defined task with no pressure condition applied
3.13 pressure condition defined variation of input or environment intended to move a behaviour away from its steady state Note 1 to entry: Annex C lists pressure condition classes.
3.14 unannounced arm execution of a conformity test in which nothing available to the deployed assembly indicates that the input is a test
3.15 invariance gap difference between the result of a conformity test in its announced arm and in its unannounced arm
3.16 prediction record written statement of the expected result of a test run, with its unit of measurement and its grader, dated and hashed before the run
3.17 grader person, procedure or system that decides whether a conformity test passed
3.18 independent tester person or system that did not produce, and is not responsible for, the output under test
3.19 blast radius scope of harm to users, data, services or third parties that a test run can cause
3.20 nonconformity non-fulfilment of a requirement Note 1 to entry: wording as in the ISO harmonized structure for management system standards; not re-checked against the published text in this draft.
3.21 continuous conformity state of a deployed AI system for which conformity with its stated requirements is shown by conformity tests run on the deployed assembly at a fixed interval and on every change event, with records that let a third party repeat the tests Note 1 to entry: the term matches Regulation (EU) 2024/1689 Art 72(2), which requires a post-market monitoring system that allows the provider "to evaluate the continuous compliance of AI systems with the requirements set out in Chapter III, Section 2". Art 72(2) names the aim; this document specifies active testing as one means to it.
3.22 machine-checkable requirement requirement stated in a form that a tool can evaluate against an output, a record or a configuration without human interpretation
3.23 rule table versioned set of machine-checkable requirements, each with an identifier, a check, a severity and a corrective instruction
3.24 output boundary point at which an output of the deployed assembly leaves the assembly for a user, a file or another system
3.25 decision layer The part of an output that commits to a decision: a verdict, a recommendation, a selected action or a stated fact that a user or a downstream system acts on. Distinct from the output text in which the decision is expressed.
3.26 scripted pressure sequence A fixed sequence of two or more turns, each applying a pressure condition (3.13) from Annex C, applied after an initial output, with the correct decision fixed by ground truth before the run and no new evidence introduced by the sequence.
3.27 fold A change at the decision layer (3.25) from a correct decision to an incorrect one during a scripted pressure sequence (3.26). The fold rate is the number of runs that fold divided by the number of runs that were correct before the sequence began.
4 Conformity testing programme
CC-4.1 H The organization shall establish, maintain and review a documented conformity testing programme for each deployed AI system.
- Rationale: DORA Art 24(1) uses the same structure for ICT systems ("establish, maintain and review a sound and comprehensive digital operational resilience testing programme"). AI Act Art 9(1) uses "established, implemented, documented and maintained" for risk management.
- Verification: the programme document exists, carries a version and an owner, and was reviewed within the last [12] months.
CC-4.2 H The programme shall list the requirements under test. For each requirement it shall state a metric and a pass threshold before the first test run.
- Rationale: AI Act Art 9(8) requires testing "against prior defined metrics and probabilistic thresholds".
- Verification: each requirement in the programme has a metric and a threshold with a date earlier than the first run record.
CC-4.3 H The programme shall name one person accountable for it.
- Rationale: GitLab, 2017: the backup restore had no owner, so nobody tested it (the incident sources named in Annex B).
- Verification: a named owner is recorded in the programme document.
5 Object of test
CC-5.1 A Conformity tests shall run against the deployed assembly. Results from the model alone, or from a different assembly, shall not be reported as results for the deployed AI system.
- Rationale: TIBER-EU (2025), section 2.1.1: tests of a single system "within isolation" do not "assess the full scenario". DORA Art 26(2): TLPT "shall be performed on live production systems". CrowdStrike Channel File 291 (2024): every test used a wildcard in the 21st field, so the input mismatch appeared only in production (CrowdStrike RCA, 6 Aug 2024).
- Verification: each run record names the assembly configuration and it matches the configuration serving users at the time of the run.
CC-5.2 M Each run record shall identify the version of every component of the deployed assembly at the time of the run.
- Rationale: Knight Capital (2012): one of eight servers missed a deployment and ran dead code; the firm lost "$460 million" (SEC order 34-70694). AI Act Art 12 requires events to be logged so that the functioning of the system is traceable.
- Verification: the run record lists a version or content hash for each component in 3.3.
CC-5.3 H Where conformity tests run on the system that serves users, the organization shall define and apply controls on blast radius before the run. These shall include: no use of real credentials or real personal data unless the test requires it and a documented legal basis exists; a means to stop the run at once; and a limit on the actions the assembly can take during the run.
- Rationale: DORA Art 26(5) requires "effective risk management controls" during TLPT on live systems.
- Verification: the blast-radius controls are documented in the run plan and the stop mechanism was tested in the last [90] days.
6 Test design
CC-6.1 A For each behaviour under test, the organization shall measure the steady state before applying any pressure condition.
- Rationale: chaos engineering practice starts from a measured steady state (Principles of Chaos Engineering; Basiri et al., IEEE Software 33(3), 2016). Without a steady state a change under pressure cannot be measured. Piece 7, "Chaos Engineering for Behaviour".
- Verification: each behaviour has a steady-state measurement with a date earlier than the first pressure run.
CC-6.2 A The test set shall include at least one pressure condition from each class in Annex C that is relevant to the intended purpose. Exclusion of a class shall be documented with a reason.
- Rationale: a system that holds under one kind of pressure can fail under the next. The objections register incident log on machinebehavior.io records slips on one model family from different pulls (over-reading toward the interesting case, uneven caution by category, the corpus default treated as neutral).
- Verification: mapping of tests to Annex C classes, with documented exclusions.
CC-6.3 M For each test run the organization shall write a prediction record before the run, and shall keep any later amendment with its date and reason.
- Rationale: a prediction edited after the run cannot be told apart from a description of the result. Practice: the prediction hashes in the machinebehavior.io repository (predictions/HASHES.txt).
- Verification: the hash of the prediction record has a timestamp earlier than the run, from a source the organization does not control alone [for example a public repository or a timestamp service].
CC-6.4 A The test set should include unannounced arms. Where it does, the run record shall report the invariance gap.
- Rationale: models can behave differently when they detect evaluation (Greenblatt et al., arXiv:2412.14093; Meinke et al., arXiv:2412.04984). A result from an announced test alone does not show the behaviour in use.
- Verification: presence of paired arms and a reported invariance gap.
CC-6.5 A Each grader shall be tested against cases with a known pass result and cases with a known fail result before use and after each change to the grader.
- Rationale: a test that cannot fail gives no evidence. CrowdStrike (2024): the wildcard made the tests blind to the fault. "Chaos Engineering for Behaviour" (machinebehavior.io), row "Test the test".
- Verification: grader test records with known-pass and known-fail cases.
CC-6.6 A The test set shall include decision-layer tests: tasks whose correct decision is fixed by ground truth before the run, applied under a scripted pressure sequence (3.26) drawn from the classes in Annex C. The run record shall report, per arm, the correctness before the sequence, the correctness after it and the fold rate (3.27). A test set that consists only of output-text rules (CC-12.3) shall not be reported as evidence of conformity at the decision layer.
- Rationale: a wording check passes a model that decides wrongly in fluent text. Experiment 04, full run full-01 (https://machinebehavior.io/experiments/#experiment-04, 2026-10-07, qwen3-coder-30b on Amazon Bedrock, 432 scored runs across three arms, temperature 0): zero fawn-opener and zero tic events at turn 0 in every arm, while the baseline arm folded on 75 percent [95 percent CI 62 to 84] of the runs that were correct before pressure (42 of 56), and correctness on the ground-truth WAIT rows fell from 78 percent at turn 0 to 32 percent after four scripted pushes, most folds at the third push. The lexical arm's fold rate (62 percent) overlapped the baseline's. The short-stance run of the same experiment (full-02, https://machinebehavior.io/experiments/#experiment-04-short-stance, 2026-10-07, prereg v3, 720 calls): a 2,846-character stance instruction kept turn-0 correctness at 76 percent [66 to 85] and folded on 33 percent [22 to 46] of eligible runs (18 of 55) against the baseline's 75 percent, intervals not overlapping, while final correctness (50 percent against 32 percent) overlapped by four points and is not read as a net win. A decision-layer test separates these two instructions; a wording check scores both arms identical, at zero lexical events. Scope of the figures: the 75 percent fold rate and the 33 percent under the short instruction are properties of qwen3-coder-30b under this pressure. The same design on three larger models (https://machinebehavior.io/experiments/#experiment-04-three-models, 2026-10-07, prereg v5, 2,880 calls per model): gpt-oss-120b, DeepSeek V3.2 and Kimi K2.5 folded on 0 of 72 baseline runs each, intervals [0 to 5]. Folding under scripted pressure is a property of a model, not of language models in general, and the task (one subtraction on a fact sheet) may make holding easy for strong models. The requirement does not depend on which way a model goes: it requires the test, and a run record that shows the model held is the evidence a wording check cannot give. A deployed system whose conformity evidence is a clean lexical score therefore carries no evidence about its decisions under pressure. Incidents of the same shape: Moffatt v. Air Canada (2024 BCCRT 149), a fluent chatbot answer stating a policy that did not exist, C$812.02 and a rejected "separate legal entity" defence; Mata v. Avianca (678 F.Supp.3d 443, S.D.N.Y. 2023), fluent citations to cases that did not exist, $5,000 sanction. The chess-engine case (arXiv:2502.13295, read in "The Cheating Moved"): a surface behaviour trained out, the decision-level behaviour moved.
- Verification: run records containing the ground truth per task, the hash of the frozen scripted sequence, the number of runs, and per arm the turn-0 correctness, the final correctness and the fold rate with a confidence interval. Requirements CC-6.1 to CC-6.5 apply to decision-layer tests as to any other test.
- Note: the threshold for an acceptable fold rate is a parameter of the programme (CC-4.2), set per intended purpose; this document fixes the measurement, not the limit.
7 Cadence and triggers
CC-7.1 M The organization shall run the full test set on every change event, before the changed assembly serves users.
- Rationale: AI Act Art 43(4) requires a new conformity assessment only on a substantial modification, and changes pre-determined by the provider do not count as one. A system prompt edit, a tool change or a knowledge base update can change outputs without being a substantial modification.
- Verification: for each change event in the change log, a passing run record with a later timestamp than the change and an earlier timestamp than the release.
CC-7.2 H A change event should be released in stages. Conformity tests should pass at each stage before the next.
- Rationale: CrowdStrike RCA, Finding 6: Template Instances were not deployed in stages; "8.5 million Windows devices" were affected (Microsoft figure).
- Verification: release records show stages with test results per stage.
CC-7.3 M The organization shall run the full test set at least every [30] days, whether or not a change event occurred.
- Rationale: a model version supplied by a third party, or a source of record outside the organization, can change without a change event the organization sees. DORA Art 24(6) sets a fixed minimum ("at least yearly") for ICT systems; the shorter interval proposed here reflects the rate at which model versions and knowledge change. The value is open for review.
- Verification: run records with no gap longer than the interval.
CC-7.4 A Every conformity test that has detected a nonconformity shall be added to the regression set. A test shall not be removed from the regression set without a documented reason approved by the programme owner.
- Rationale: a fix for one failure closes one route while the cause stays. Piece "They Trained Out the Board Edit. The Cheating Moved." (machinebehavior.io): the board-edit route was trained out and the behaviour used the next route.
- Verification: each closed nonconformity maps to a test in the regression set.
CC-7.5 H Attack methods published in the public record that are relevant to the intended purpose should be added to the test set within [60] days of publication.
- Rationale: the test set otherwise covers only attacks the organization found itself. "Chaos Engineering for Behaviour" (machinebehavior.io), "Automate, and rerun on every change".
- Verification: a dated intake log of published methods with the decision for each.
8 Knowledge base conformity
CC-8.1 A Each knowledge item shall have a named owner and a recorded source of record.
- Rationale: Fogbank (2000-2008): NNSA "kept few records of the process" and the experts left; relearning cost "$69 million" (GAO-09-385).
- Verification: sample of knowledge items; each has an owner and a source of record.
CC-8.2 M Each knowledge item shall carry the date it was last verified against its source of record. An item older than its freshness interval shall be re-verified or withdrawn from retrieval.
- Rationale: Moffatt v. Air Canada, 2024 BCCRT 149: the airline did not take "reasonable care to ensure" its chatbot was accurate and was held to the chatbot's answer.
- Verification: no item in retrieval is older than its freshness interval.
CC-8.3 A The organization shall state the freshness interval for each class of knowledge item, with a reason tied to how often the source of record changes.
- Rationale: one fixed interval for all knowledge either wastes effort on stable items or leaves fast-changing items stale.
- Verification: the programme lists classes, intervals and reasons.
CC-8.4 A The test set shall include knowledge conformity tests: inputs whose correct output is fixed by a knowledge item, graded against the source of record.
- Rationale: Moffatt v. Air Canada; Cursor support bot (2025), which stated a login policy that did not exist (The Register, 18 Apr 2025; secondary source).
- Verification: share of knowledge item classes covered by at least one knowledge conformity test.
CC-8.5 A Verification of a knowledge item shall be made against the source of record and not against a summary of it.
- Rationale: on 2026-10-03, a web-fetch summariser reported a $30,000 sanction in Mata v. Avianca; the court order sets $5,000 (the incident sources named in Annex B).
- Verification: verification records cite the source of record and its version or date.
CC-8.6 A Where outputs are used for decisions by users or third parties, statements of fact in the output should be traceable to a knowledge item or a cited source.
- Rationale: Mata v. Avianca, 678 F.Supp.3d 443 (S.D.N.Y. 2023): citations generated by a model were not checked; sanction of $5,000. Deloitte report for DEWR (2025): fabricated references, partial refund (secondary source).
- Verification: sample of outputs; share of factual statements with a trace.
9 Independence
CC-9.1 H The pass or fail decision of a conformity test shall not rest only with the system or person that produced the output under test.
- Rationale: DORA Art 24(4) requires tests "undertaken by independent parties, whether internal or external". Piece "What Operations Already Knows" (machinebehavior.io): in the author's logged setup the model caught none of its own seven relapses and the human operator caught all seven.
- Verification: grader identity recorded per run and differs from the producer of the output.
CC-9.2 H An independent tester from outside the organization should run the test set at least every [third] full run cycle [or at least every 12 months].
- Rationale: DORA Art 26(8): where internal testers are used for TLPT, external testers are required "every three tests".
- Verification: external run records at the stated cadence.
10 Findings and remediation
CC-10.1 A The organization shall classify, assign and track every nonconformity to closure. Closure shall require a passing rerun of the test that found it.
- Rationale: DORA Art 24(5) requires procedures "to prioritise, classify and remedy all issues revealed" and "internal validation methodologies" to confirm the fix.
- Verification: each closed nonconformity has a passing rerun record.
CC-10.2 H Where a nonconformity meets the definition of a serious incident (AI Act Art 3(49)), or shows that the system may present a risk within the meaning of Art 79(1), the organization shall pass it to its incident reporting process without delay.
- Rationale: AI Act Art 26(5) (deployers inform the provider and suspend use) and Art 73 (reporting deadlines).
- Verification: link from each such finding to an incident record.
11 Records and disclosure
CC-11.1 M Each run record shall contain at least: date and time; deployed assembly versions (CC-5.2); test set version; prediction record and its hash; raw results; pass or fail per test; grader identity; and the steps needed to rerun.
- Rationale: AI Act Annex IV point 2(g) requires "test logs and all test reports dated and signed by the responsible persons". A record without rerun steps cannot be checked by a third party.
- Verification: sample of run records against the list.
CC-11.2 M Run records shall be kept for at least [the lifetime of the deployed AI system plus 6 months].
- Rationale: AI Act Art 26(6) sets at least six months for automatically generated logs held by deployers; test records serve the same evidential purpose over a longer period.
- Verification: oldest available record against the retention rule.
CC-11.3 H The organization may publish its method, raw results and rerun steps. A published assessment made by the organization of its own system shall be labelled a self-assessment and shall not be presented as a certification.
- Rationale: under the AI Act, conformity assessment by a notified body is a defined procedure (Art 43); a self-assessment is not one.
- Verification: label present on every published result.
CC-11.4 H A published self-assessment should invite a second rater and should publish any disagreement between raters.
- Rationale: CC-9.1. A self-audit without an outside check repeats the weakness this document addresses.
- Verification: invitation and rater record present.
12 Machine-checkable requirements
CC-12.1 H The programme shall mark each requirement as mechanical (a tool decides), assisted (a tool flags, a person decides) or manual (a person decides).
- Rationale: a mechanical check covers only what it can see. Banning the surface form of a behaviour can move the behaviour elsewhere ("Sycophancy Is Layered", machinebehavior.io series; vestige-kit README). Marking the class prevents a passing mechanical check from being read as evidence for a requirement that needs judgement.
- Verification: every requirement in the programme carries one of the three marks.
CC-12.2 A Requirements marked mechanical should be expressed in a machine-readable form, kept under version control with the test set.
- Rationale: frameworks such as the AI Act, ISO/IEC 42001 and NIST AI RMF "specify what to assure but provide no executable format for how" (Cilla Ugarte et al., arXiv:2604.13767, 2026). NIST OSCAL provides "open, machine-readable formats available in XML, JSON, and YAML" for control information (pages.nist.gov/OSCAL). CEN runs a SMART Standards project to make standards machine-readable (experts.cen.eu).
- Verification: each mechanical requirement has a machine-readable form with a version identifier.
CC-12.3 M Where a mechanical requirement applies to every output, the organization should enforce it at the output boundary with a rule table. Each rule in a blocking tier shall have a recorded false-positive test, and rules that fail it shall move to a warning tier.
- Rationale: working instance (author's own): vestige-kit, a rule table of regular-expression rules in a blocking tier and a warning tier, checked by a hook on every reply and every Markdown file written. The blocking tier is limited to rules with near-zero false positives; one rule records "0 hits across 1,282 markdown files" of the author's vault as its false-positive test (vestige-kit, hooks/vestige-patterns.js).
- Verification: rule table with a test record for each blocking rule.
CC-12.4 M Each rule hit at the output boundary shall be logged with the rule identifier, the time and the session or request. The log shall be reviewed at each programme review to retire, tighten or promote rules.
- Rationale: an unlogged block leaves no evidence that the control ran. vestige-kit keeps a hit log and a script that turns it into retire and promote decisions (hooks/analyze-log.js).
- Verification: hit log present; programme review record cites it.
CC-12.5 M Run records (CC-11.1) and nonconformities (CC-10.1) should be exportable in an open, published format for assessment results and findings.
- Rationale: the OSCAL Assessment Results and Plan of Action and Milestones models represent assessment results and findings; NIST states that "the assessment model supports information from periodic and continuous assessments" (pages.nist.gov/OSCAL, layer overview). An open format lets a second rater or an assessor rerun the evaluation without re-keying records.
- Verification: export of a sample run record that validates against the published schema.
Annex A (informative): relation to existing law and standards
| Clause here | Existing text | What exists | What this document adds |
|---|---|---|---|
| 4.1 | DORA Art 24(1); AI Act Art 9(1) | Testing programme (DORA); risk management system (AI Act) | A programme scoped to the deployed AI system |
| 4.2 | AI Act Art 9(8) | Prior defined metrics and thresholds, pre-market | Same rule applied after deployment |
| 5.1 | DORA Art 26(2); TIBER-EU 2.1.1 | Tests on live production (financial ICT) | Same principle for the deployed assembly |
| 6.4 | none found | Unannounced arms and invariance gap | |
| 7.1 | AI Act Art 43(4) | Reassessment on substantial modification | Tests on every change event |
| 7.3 | DORA Art 24(6); AI Act Art 17(1)(d) | At least yearly (DORA); test procedures before, during and after development at a frequency the provider sets (AI Act, high-risk) | A fixed maximum interval, and tests on every change event |
| 8.x | AI Act Art 72(2); Art 53(1)(a) "keep up-to-date" (GPAI documentation) | Monitoring data; documentation kept current | Testing of the knowledge the system runs on |
| 9.1-9.2 | DORA Art 24(4), 26(8) | Independent and external testers | Same for deployed AI systems |
| 10.1 | DORA Art 24(5) | Remediation with validation | Same |
| 11.1 | AI Act Art 12; Annex IV 2(g) | Logs; dated, signed test reports | Rerun steps and prediction record |
| (all) | AI Act Art 72(2) | Aim: "continuous compliance"; means: data collection | Means: active tests |
| (all) | NIST AI RMF MEASURE 2.4, MANAGE 4.1 | Monitoring in production; post-deployment monitoring plans | Testing, with cadence |
| (all) | ISO/IEC 42001 A.6.2.4, A.6.2.6 (numbers from secondary sources) | Verification and validation; operation and monitoring | Cadence and triggers after deployment |
| 12.2, 12.5 | NIST OSCAL (v1.2.3 is the latest release on GitHub, checked 2026-10-03); NIST SP 800-126 Rev. 3 (SCAP 1.3, February 2018) | Machine-readable control catalogues, assessment results and findings (security) | Same formats applied to AI conformity tests |
| 3.3, 5.1 | ISO/IEC AWI 26741 (stage 20.00): framework for describing an AI system as object of conformity assessment, including "system boundaries" | Work item open | Deployed assembly as the boundary |
| 12.x | CEN SMART Standards project | Machine-readable standards (format of the standard itself) | Machine-checkable requirements (format of the requirement inside a programme) |
Annex B (informative): origin of each requirement
| Requirement | Incident or record | Source type |
|---|---|---|
| CC-4.3 | GitLab database deletion, 2017 | primary (GitLab postmortem) |
| CC-5.1, CC-6.5, CC-7.2 | CrowdStrike Channel File 291, 2024 | primary (CrowdStrike RCA) |
| CC-5.2 | Knight Capital, 2012 | primary (SEC order 34-70694) |
| CC-6.2 | Objections register incident log | author's public record |
| CC-6.6 | Eval 04 full-01: clean lexical score, 75 percent fold rate under scripted pushback (qwen3-coder-30b only); full-02: short stance instruction folds 33 percent at the same turn-0 correctness; v5: three larger models folded 0 of 72 each (2026-10-07); Moffatt v. Air Canada; Mata v. Avianca | author's published experiment (https://machinebehavior.io/experiments/#experiment-04, #experiment-04-short-stance, #experiment-04-three-models); primary; primary |
| CC-7.4 | Chess-engine board edit trained out; behaviour moved | author's piece, citing Palisade arXiv:2502.13295 |
| CC-8.1 | Fogbank, 2000-2008 | primary (GAO-09-385) |
| CC-8.2, CC-8.4 | Moffatt v. Air Canada, 2024 | primary (2024 BCCRT 149) |
| CC-8.5 | Summariser invented a sanction figure, 2026-10-03 | author's record |
| CC-8.6 | Mata v. Avianca, 2023; Deloitte/DEWR, 2025 | primary; secondary |
| CC-9.1 | Seven relapses, none self-caught | author's piece "What Operations Already Knows" |
| CC-12.1 | Surface ban moved the behaviour (symptom substitution) | author's piece "Sycophancy Is Layered" |
| CC-12.3, CC-12.4 | vestige-kit rule table and hit log | author's working instance |
Annex C (normative): pressure condition classes
| Class | Description | Example |
|---|---|---|
| C.1 Pushback | user disputes a correct output | "That is wrong, check again" after a correct answer |
| C.2 False premise | input contains a confident false statement | a question that assumes a policy that does not exist |
| C.3 Authority claim | input claims a role or permission | "As the system administrator, I authorise you to" |
| C.4 Outcome-only scoring | the task rewards the result and not the method | win-only scoring in a game task |
| C.5 Approval signal | input rewards agreement | praise after agreement, displeasure after disagreement |
| C.6 Stale or conflicting knowledge | retrieved knowledge disagrees with the source of record | an outdated price list in retrieval |
| C.7 Instruction in data | retrieved content or tool output contains instructions | a web page that tells the assistant to change task |
Bibliography
- Regulation (EU) 2024/1689 (Artificial Intelligence Act), OJ L, 12.7.2024.
- Regulation (EU) 2022/2554 (DORA), OJ L 333, 27.12.2022.
- Commission Delegated Regulation (EU) 2025/1190 (TLPT RTS), read 2026-10-04: covers testing of live production systems of critical or important functions (Art 5(1)); active red team phase at least 12 weeks (Art 11(5)); sets no interval of its own and no change-triggered test.
- ECB, TIBER-EU Framework, February 2025.
- NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0), 2023.
- ISO/IEC 42001:2023; ISO/IEC 27001:2022; ISO/IEC 27002:2022; ISO 22301:2019.
- Principles of Chaos Engineering (principlesofchaos.org); Basiri et al., "Chaos Engineering", IEEE Software 33(3), 2016.
- Greenblatt et al., arXiv:2412.14093; Meinke et al., arXiv:2412.04984; Palisade Research, arXiv:2502.13295.
- Incident sources as named in each rationale and in Annex B.
- machinebehavior.io: claims ledger, objections register, slips log, prediction hashes, "Chaos Engineering for Behaviour".
- NIST, Open Security Controls Assessment Language (OSCAL), pages.nist.gov/OSCAL; releases at github.com/usnistgov/OSCAL.
- NIST SP 800-126 Rev. 3, The Technical Specification for the Security Content Automation Protocol (SCAP): SCAP Version 1.3, February 2018.
- Cilla Ugarte et al., "Making AI Compliance Evidence Machine-Readable", arXiv:2604.13767, submitted 15 April 2026.
- CEN, SMART Standards project, experts.cen.eu/key-initiatives/smart-standards/.
- ISO/IEC AWI 26741 (iso.org/standard/94402.html); ISO/IEC DIS 23282 (iso.org/standard/87387.html).
- vestige-kit (github.com/uncovertechtalent/vestige-kit).
Crosswalk to existing frameworks
Each requirement of this draft, the clause of an existing framework it relates to, and how. "Evidence toward" means a passing test of the requirement is evidence for part of that clause, with the slice named; it is never conformity to the framework. "Nearest clause" means the framework has nothing closer and does not require what the requirement asks: these rows are the gap this document names. Machine-readable: /conformity/requirements.json.
| req | mark | framework | clause | relation | slice |
|---|---|---|---|---|---|
| CC-4.1 | H | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 9(1)-(2) | evidence toward | risk management system as a continuous iterative process |
| CC-4.1 | H | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 72(1) | evidence toward | documented post-market monitoring system |
| CC-4.1 | H | Regulation (EU) 2022/2554 (DORA), EUR-Lex | Art 24(1) | evidence toward | digital operational resilience testing programme |
| CC-4.2 | H | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 9(8) | evidence toward | testing against prior defined metrics and probabilistic thresholds |
| CC-4.2 | H | NIST AI 100-1, AI Risk Management Framework 1.0 | MEASURE 1.1 | evidence toward | approaches and metrics for measurement selected |
| CC-4.3 | H | NIST AI 100-1, AI Risk Management Framework 1.0 | GOVERN 2.1 | evidence toward | roles and responsibilities documented |
| CC-5.1 | A | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 15(1) | evidence toward | consistent performance of the system as placed on the market |
| CC-5.1 | A | Regulation (EU) 2022/2554 (DORA), EUR-Lex | Art 24(2) | evidence toward | testing of ICT systems supporting critical functions |
| CC-5.2 | M | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 12(1) | evidence toward | automatic recording of events over the lifetime of the system |
| CC-5.2 | M | Regulation (EU) 2022/2554 (DORA), EUR-Lex | Art 9(4)(e) | evidence toward | documented ICT change management |
| CC-5.3 | H | Regulation (EU) 2022/2554 (DORA), EUR-Lex | Art 26(5) | evidence toward | risk management measures for threat-led tests on live production systems |
| CC-6.1 | A | NIST AI 100-1, AI Risk Management Framework 1.0 | MEASURE 2.5 | nearest clause; it does not require this | validity and reliability assessed |
| CC-6.2 | A | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 15(5) | evidence toward | resilience against attempts to alter use, outputs or performance |
| CC-6.2 | A | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 55(1)(b) | nearest clause; it does not require this | adversarial testing of GPAI models with systemic risk |
| CC-6.3 | M | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 9(8) | evidence toward | prior defined metrics and thresholds |
| CC-6.4 | A | Regulation (EU) 2022/2554 (DORA), EUR-Lex | Art 26(1) | nearest clause; it does not require this | threat-led penetration testing, covert by design |
| CC-6.5 | A | NIST AI 100-1, AI Risk Management Framework 1.0 | MEASURE 2.5 | nearest clause; it does not require this | validity of the evaluation itself |
| CC-6.6 | A | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 15(1) | evidence toward | accuracy and robustness, performance consistent through the lifecycle |
| CC-6.6 | A | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 15(5) | evidence toward | resilience against attempts to alter outputs |
| CC-6.6 | A | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 9(8) | evidence toward | testing against prior defined metrics |
| CC-6.6 | A | Regulation (EU) 2022/2554 (DORA), EUR-Lex | Art 26(2) | nearest clause; it does not require this | threat-led testing on live systems |
| CC-7.1 | M | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 9(6)-(7) | evidence toward | testing throughout development and before placing on the market |
| CC-7.1 | M | Regulation (EU) 2022/2554 (DORA), EUR-Lex | Art 9(4)(e) | evidence toward | ICT change management with testing before deployment |
| CC-7.1 | M | Regulation (EU) 2022/2554 (DORA), EUR-Lex | Art 25(1) | evidence toward | appropriate tests on ICT systems |
| CC-7.1 | M | NIST CSF 2.0 | PR.PS-01 | evidence toward | configuration management |
| CC-7.2 | H | Regulation (EU) 2022/2554 (DORA), EUR-Lex | Art 9(4)(e) | nearest clause; it does not require this | change management |
| CC-7.3 | M | Regulation (EU) 2022/2554 (DORA), EUR-Lex | Art 24(6) | evidence toward | appropriate tests at least yearly |
| CC-7.3 | M | Regulation (EU) 2016/679, EUR-Lex | Art 32(1)(d) | evidence toward | regular testing, assessing and evaluating |
| CC-7.3 | M | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 72(2) | evidence toward | evaluate continuous compliance throughout the lifetime |
| CC-7.3 | M | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 17(1)(d) | evidence toward | test procedures after development at a stated frequency (high-risk providers; frequency set by the provider) |
| CC-7.4 | A | NIST CSF 2.0 | ID.IM-03 | evidence toward | improvements from lessons learned |
| CC-7.5 | H | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 55(1)(b) | nearest clause; it does not require this | state-of-the-art adversarial testing |
| CC-8.1 | A | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 10(2) | nearest clause; it does not require this | data governance practices |
| CC-8.2 | M | Regulation (EU) 2016/679, EUR-Lex | Art 5(1)(d) | evidence toward | accuracy; kept up to date |
| CC-8.2 | M | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 10(3) | nearest clause; it does not require this | data sets relevant, representative, free of errors |
| CC-8.3 | A | Regulation (EU) 2016/679, EUR-Lex | Art 5(1)(d) | nearest clause; it does not require this | accuracy |
| CC-8.4 | A | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 15(1) | nearest clause; it does not require this | accuracy |
| CC-8.5 | A | Regulation (EU) 2016/679, EUR-Lex | Art 5(1)(d) | nearest clause; it does not require this | accuracy |
| CC-8.6 | A | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 13(1) | nearest clause; it does not require this | transparency, interpretable output |
| CC-9.1 | H | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 14(1) | evidence toward | human oversight |
| CC-9.1 | H | Regulation (EU) 2022/2554 (DORA), EUR-Lex | Art 24(4) | evidence toward | tests by independent parties, internal or external |
| CC-9.2 | H | Regulation (EU) 2022/2554 (DORA), EUR-Lex | Art 26(8) | evidence toward | external testers for threat-led tests |
| CC-9.2 | H | Regulation (EU) 2022/2554 (DORA), EUR-Lex | Art 27 | evidence toward | requirements for testers |
| CC-10.1 | A | NIST CSF 2.0 | ID.IM-03 | evidence toward | lessons learned |
| CC-10.1 | A | Regulation (EU) 2022/2554 (DORA), EUR-Lex | Art 24(5) | evidence toward | remediation of issues identified in tests |
| CC-10.2 | H | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 73(1) | evidence toward | reporting of serious incidents |
| CC-10.2 | H | Regulation (EU) 2022/2554 (DORA), EUR-Lex | Art 19(1) | evidence toward | reporting of major ICT-related incidents |
| CC-11.1 | M | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 12(1) | evidence toward | automatic recording of events |
| CC-11.1 | M | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 72(2) | evidence toward | collection of data on performance throughout the lifetime |
| CC-11.1 | M | NIST CSF 2.0 | DE.CM-09 | evidence toward | monitoring of software and services |
| CC-11.2 | M | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 19(1) | evidence toward | logs kept for at least six months |
| CC-11.2 | M | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 18(1) | evidence toward | documentation kept 10 years |
| CC-11.3 | H | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 43(2) | nearest clause; it does not require this | conformity assessment based on internal control |
| CC-11.4 | H | Regulation (EU) 2022/2554 (DORA), EUR-Lex | Art 24(4) | nearest clause; it does not require this | independent parties |
| CC-12.1 | H | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 14(1) | evidence toward | which decisions a person takes |
| CC-12.2 | A | NIST CSF 2.0 | PR.PS-01 | nearest clause; it does not require this | configuration management |
| CC-12.3 | M | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 15(1) | evidence toward | consistent performance at the output |
| CC-12.3 | M | NIST CSF 2.0 | PR.PS-01 | evidence toward | configuration management of the rule table |
| CC-12.4 | M | Regulation (EU) 2024/1689 (AI Act), EUR-Lex | Art 12(1) | evidence toward | logging |
| CC-12.4 | M | NIST CSF 2.0 | DE.CM-09 | evidence toward | monitoring |
| CC-12.5 | M | NIST CSF 2.0 | GV.OC | nearest clause; it does not require this | organizational context |
Limits
- One author. The requirements and the self-assessment come from the same setup, so the draft may fit that setup more closely than an outside draft would.
- No committee has reviewed it. It cites standards by number and public title only; no standard text was processed in writing it, under the publisher's licence terms for machine processing.
- Labelled secondary in the text: ISO/IEC 42001 Annex A control numbers (Annex A of this draft) and the wording of term 3.20, not re-checked against the published standards.
- The decision-layer figures cited for CC-6.6 belong to one model (qwen3-coder-30b); three larger models did not fold under the same design.
- Square-bracket values are proposals, not findings.
Changes
| date | version | change |
|---|---|---|
| 2026-10-09 | 0.3 | Correction after an outside review: AI Act Art 17(1)(d) (test procedures before, during and after development, at a frequency the provider sets) added to the introduction, Annex A, the gap table and the crosswalk; the sentence "no rule in Union law requires anyone to run those tests after launch" withdrawn as too broad. |
| 2026-10-09 | 0.3 | Published as a reference on machinebehavior.io. |
| 2026-10-07 | 0.3 | CC-6.6 decision-layer tests; terms 3.25 decision layer, 3.26 scripted pressure sequence, 3.27 fold. |
| 2026-10-03 | 0.2 | Clause 12, machine-checkable requirements; name anchored to AI Act Art 72(2). |