IRBAI · Audit & Oversight

How We Audit AI Systems

Certification states what a system is permitted to do; audit verifies what it actually does, so that AI systems remain safe for the public for as long as they operate. IRBAI conducts audits directly and publishes the frameworks and guidelines under which accredited third parties conduct them worldwide. IRBAI audits are conducted under Chapter 16 of the Statute (Articles 115–121) to the technical standards of Annex II, by IRBAI auditors and by Accredited Assessment Bodies under Article 102. Certified entities accept and facilitate audit as a condition of certification. In return, audit access is bounded, logged, and confidential by design.

How We Regulate: the framework How We Certify: the assessment
01 · VERIFY, DON’T TRUST

An audit tests certified claims against deployed reality. Documentation is evidence, not proof. The standard of verification is what a system does in operation: its logs, and its behaviour under adversarial pressure.

02 · PROVE WITHOUT DISCLOSURE

Wherever possible, compliance is proven cryptographically, with no access to raw materials. Where access is unavoidable, it occurs through controlled mechanisms from which raw weights and training data are never exported.

03 · CONTINUOUS, NOT ANNUAL

Audit follows the certificate through its life: scheduled renewal, six-month adversarial testing for Systemic systems, for-cause audits at any time, and mandatory reassessment on material change.

SCOPE · THE EIGHT AUDIT ACCESS OBJECTS (ANNEX II, PART 1)

Within a lawfully initiated audit under Article 116, auditors can access eight categories of material, covering a system from silicon to deployed operation. Access is scoped to the certificate, purpose-bound, and logged.

Training data & lineage
Datasets, provenance, licensing, augmentation histories.
Part 1(a)
Pipelines & algorithms
Preprocessing, training algorithms, optimization procedures.
Part 1(b)
Model artifacts
Checkpoints, version manifests, architecture documentation.
Part 1(c)
Evaluation & safety records
Methodologies, benchmark results, safety assessments.
Part 1(d)
Hardware & physical systems
Hardware, IoT integrations, robotics interfaces.
Part 1(e)
Telemetry & logs
Deployment telemetry; append-only operational audit logs.
Part 1(f)
Security & incidents
Security testing, vulnerability assessments, incident response.
Part 1(g)
Supply chain
Dependencies, pre-trained components, external APIs.
Part 1(h); Article 111

METHOD · THE VERIFICATION LADDER

Annex II sets a layered model: prove cryptographically where possible, access confidentially where necessary, test adversarially throughout. Each rung is used only where the previous one cannot achieve verification.

First resort: no access to raw materialsRung 01
Cryptographic verification
Compliance is proven mathematically (Annex II Part 3): hashes and Merkle commitments fix the identity of registered models and datasets; remote attestation proves the deployed system is the certified artifact, byte for byte; zero-knowledge proofs demonstrate that training data meets licensing, consent, and provenance requirements without disclosing the data itself; software bills of materials attest the supply chain.
Annex II Part 3 · Article 119
Where cryptography is insufficientRung 02
Confidential access
Assessment proceeds through one of three mechanisms, selected in consultation with the developer (Annex II Part 4): on-premises, where auditors run approved test suites at the developer’s own facilities and remove nothing; secure remote, where an attested confidential-computing environment returns verifiable results without exposing weights; or accredited third-party assessment within an attested environment. In every mechanism, raw weights and training data are never transferred to IRBAI. Escrow safeguards and the Break-Glass Emergency Access Protocol are governed exclusively by the Statute (Article 119).
Annex II Part 4 · Article 119
Applied throughout, at the elicitation standardRung 03
Independent adversarial testing
For High Risk and Systemic AI Systems, testing covers at minimum: prompt injection (direct, indirect, multi-turn); data exfiltration, including training-data and personal-data extraction; model exploitation (extraction, inversion, adversarial perturbation); and safety guardrail circumvention, including jailbreaking, refusal suppression, and capability elicitation beyond certified parameters. Methodologies are maintained by the Compliance Monitoring Directorate and updated as attack techniques evolve.
Annex II Part 2 · Article 116

IN PRACTICE · THE AUDIT ENGAGEMENT

Inside that method, an audit engagement follows four familiar activities. Each is conducted to the Annex II standards and scoped to the certificate under review.

Risk assessment
The system is examined for risk-relevant behaviour: bias, security vulnerabilities, and safety concerns, assessed against its tier and its position in the risk matrix.
Annex I Part 1 · Article 116
Documentation review
The adequacy of the Annex II artifact set is examined: development and deployment documentation, evaluation records, and monitoring practices.
Annex II Parts 1 and 3
On-site inspection
Auditors observe the system in operation and assess real-world compliance, working at the developer’s facilities through the confidential access mechanisms.
Annex II Part 4
Stakeholder interviews
Auditors engage directly with developers, operators, and affected parties, so that findings reflect how the system behaves for the people it touches.
Article 116 · Chapter 16

WHO AUDITS · AND WHO IS OBLIGED

Independence by construction

IRBAI auditors and Accredited Assessment Bodies

Audits are conducted by IRBAI and by Accredited Assessment Bodies accredited under Article 102: independent organisations operating within defined scopes of accreditation, under IRBAI oversight. No body audits a system in whose development it participated; accreditation is suspended for breach of confidentiality or misuse of audit instruments. In the Foundation Model class, evaluators additionally hold AIFM-EVAL certification.

The obligation runs beyond developers: computing infrastructure providers serving Systemic and High Risk certified systems maintain records sufficient to verify systemic-designation thresholds, with access controls enabling post-hoc attribution of compute usage (Annex II Part 5; Article 120). Developers of systems subject to certification under Chapter 14, above all High Risk and Systemic systems, register metadata in the Global Model and Dataset Registry: architecture category, approximate training compute, data categories, deployment footprint, and known limitations. Routine model updates do not touch the Registry; entries are refreshed at each certification renewal, and earlier only where a registered parameter materially changes (Annex II Part 6; Article 118).

CADENCE · WHEN AUDITS OCCUR

A certificate does not remain valid on its own. It is subject to periodic reassessment at the intervals fixed by the Statute, and to reassessment outside those intervals where the system or its circumstances change.

Every 12 months
Maximum Tier C certificate validity; renewal requires reassessment against current standards (Article 101).
Every 6 months
Independent adversarial testing for Systemic AI Systems (Article 103).
For cause, at any time
Audits lawfully initiated under Article 116; reclassification review under Article 97 where findings warrant.
On material change, within 24h
Capability breakthroughs, behavioural emergence, and material modifications: notification within twenty-four hours, then reassessment (Article 103).

FOR AUDITED ENTITIES · WHAT TO EXPECT

What triggers an audit?

Scheduled renewal, the Systemic testing cycle, material change to the system, and for-cause initiation under Article 116, including audit-log signals, incident reports, and credible third-party reports of non-compliance.

What happens to our confidential IP?

Raw weights and training data never leave your control: verification is cryptographic first, and any required access runs through the Annex II Part 4 mechanisms inside attested environments. Access is logged, purpose-bound, and protected by the confidentiality provisions of Chapter 16.

What if findings go against us?

Findings ground proportionate consequences on the Article 138 ladder (certificate conditions, suspension, and beyond) and administrative penalties within the Article 106 ceilings. A material gap between submitted and independently measured capability is itself a certification finding and may trigger reclassification review (Article 97).

Can we dispute findings?

Yes. Every adverse decision carries the Article 144A guarantees of written notice, access to the evidence relied upon, opportunity to respond, and a reasoned decision, and is appealable to the independent Appeals Panel under Article 43.

Are pre-release systems audited differently?

Yes. Before deployment, assessment is the certification procedure of Chapter 14: design, capability, and safeguards, measured at the elicitation ceiling. Once deployed, Chapter 16 audit adds operational evidence: append-only logs, incident and near-miss records, and the handling of certificate conditions.

Who conducts the audit itself?

IRBAI auditors, or an independent Accredited Assessment Body accredited under Article 102 and operating under IRBAI oversight. Accredited bodies conduct inspections, documentation reviews, and interviews, and their findings carry the same standing; a body never audits a system it helped develop.

What if we fail before release?

Certification is withheld for the assessed configuration until the findings are remediated. IRBAI defines the corrective actions and the re-assessment procedure, and escalation de-escalates upon remediation (Article 138); a restricted configuration may be separately assessed while remediation proceeds.

How do we prepare?

Maintain the Annex II artifact set continuously (documentation, append-only logs, cryptographic commitments, SBOMs, current Registry entries) and run internal evaluations to the same adversarial standard. An audit then verifies what already exists, rather than reconstructing it.

Instruments. This page summarises Chapter 16 of the IRBAI Statute (Articles 115–121) and Annex II (Audit and Compute Oversight: Technical Standards); in any divergence the Statute and Annex II prevail. The operational protocols implementing Annex II Part 1 are published as the IRBAI Audit Protocol Series. See also How We Regulate and How We Certify.