IRBAI · Certification Procedure

How We Certify AI Models

Certification is the procedure by which IRBAI assesses a model against the risk matrix and determines whether, and on what conditions, it may be certified for deployment. It is conducted under Chapter 14 of the Statute (Articles 100–109), against the mandatory assessment criteria of Annex I, and to the independent adversarial-testing standards of Annex II. The procedure produces a certification determination — a decision under Chapter 14. It does not set a system’s risk tier: the four tiers are fixed by Article 95 of the Statute.

This page describes the procedure. For the framework it applies — where regulatory risk sits — see How We Regulate.
01 · ASSESSED AT ELICITATION

Capability is measured at the level a resourced and determined actor could elicit — not the level a system displays under default use. A measurement taken below that ceiling understates risk and is not accepted as evidence.

02 · EVIDENCE BEFORE ASSERTION

Independent evaluation by an Accredited Assessment Body takes precedence over a developer’s own claims. A material gap between submitted and independently measured capability is itself a certification finding.

03 · A DETERMINATION, NOT A SCORE

Assessment yields one of four certification outcomes. These are decisions under Chapter 14; they do not alter a system’s statutory classification, and they are never averaged into a single number.

WHAT WE MEASURE — FOUR QUANTITIES, EACH EVIDENCING A STATUTORY CRITERION

Annex I fixes the mandatory criteria every system is assessed against. For foundation models, IRBAI gives those criteria evidential content through four measured quantities. Each is evidence toward a statutory criterion, not a substitute for it.

Domain capability
Performance on the capability battery defined for the risk cell — the tasks in which that cell’s risk arises.
Health & safety, economic, societal, and peace-and-security criteria (Annex I Part 1)
Generality
Breadth of competence across tasks and across cells — the degree to which a system approaches general-purpose capability.
General-purpose / AGI assessment (Article 103; Charter I)
Autonomy
Success on long-horizon, multi-step tasks under declining human oversight, including the capacity to act without a human authorisation gate.
Human Oversight & Control; Containment Integrity (Article 174A)
Marginal risk
The uplift a system provides toward serious harm, relative to what is already publicly achievable. Governs every access and release decision.
Irreversibility & Catastrophic Risk criterion (Annex I Part 1)

THE ELICITATION STANDARD

Measured at the ceiling, not the default

Risk is assessed at the capability a determined actor could draw out

All capability and marginal-risk measurements are made at the level a resourced adversary could elicit. This operationalises Annex II Part 2, under which independent adversarial testing must cover, at minimum, prompt injection, data exfiltration, model exploitation, and “safety guardrail circumvention, including jailbreaking, refusal suppression, and capability elicitation beyond certified parameters.”

The assessment adversary is calibrated to the access mode — adversarial prompting; scaffolding and orchestration; tool and retrieval access; sampling at scale; and, wherever weights are or may become accessible, fine-tuning to remove or degrade safety behaviour. For any open-weight release, the ceiling always includes fine-tuning, because released weights place safety training within the recipient’s control.

THE EVIDENCE — IN FIXED ORDER OF PRECEDENCE

Basis of the determinationPrecedence 01
Independent evaluation
Assessment against IRBAI’s private evaluation batteries, conducted by IRBAI or by an Accredited Assessment Body under Article 102. Where confidential materials are required, evaluation proceeds through the Annex II Part 4 access mechanisms — on-premises, secure remote to an attested environment, or accredited third-party — so that raw weights and training data are never exported to IRBAI. Batteries are held privately, rotated on schedule, and never disclosed to applicants.
Weighed against the independent measurePrecedence 02
Developer submission
The applicant’s Systemic Risk Assessment (Article 103), capability reporting, model documentation, red-team results, the applicable safety framework and its thresholds, and the cryptographic verification artifacts required by Annex II Part 3. Developer evidence is weighed only where its methodology meets the elicitation standard.
For renewal and deployed systemsPrecedence 03
Operational evidence
Append-only audit logs, incident and near-miss records, misuse-detection performance, and the handling of prior certification conditions — the running record of a system in service.

THE DETERMINATION — FOUR CERTIFICATION OUTCOMES

Assessment against a cell yields one of four outcomes. Where measurements span outcomes across the four quantities, the most restrictive outcome reached on any decision-relevant indicator governs; quantities are not averaged.

Certify

No material risk-relevant capability beyond the public frontier, or capability fully bounded by design.

Standard certificate conditions and twelve-month renewal (Article 101).

Certify with conditions

Measurable risk-relevant capability, with mitigations demonstrated effective at the elicitation ceiling.

Cell-specific conditions under Article 104 — access gating, monitoring, disclosure, deployment limits — and a shortened renewal interval.

Withhold pending mitigation

Material capability whose mitigations fail, or cannot be demonstrated, at the ceiling.

Certification withheld for the assessed configuration; a more restricted configuration may be separately assessed.

Refer as Prohibited

A capability falling within a Charter II category that cannot be reliably isolated or disabled (Article 95(b)).

Categorically non-certifiable under Article 95 Tier D — no condition can cure it. Referred to the Legal Committee and to enforcement under Chapter 20.

RENEWAL AND REASSESSMENT

Certification is not permanent. High Risk (Tier C) certification is valid for no more than twelve months (Article 101); Systemic AI Systems additionally undergo independent adversarial testing at intervals not exceeding six months (Article 103). Reassessment is required earlier on any capability breakthrough, unexpected behavioural emergence, or material modification — with mandatory notification within twenty-four hours. A change in access mode, in particular any move toward open-weight release, and any containment breach are in themselves reassessment triggers.

Every 12 months
Standard Tier C renewal against the current battery (Article 101).
Every 6 months
Additional adversarial testing for Systemic AI Systems (Article 103).
On material change
Immediate reassessment; 24-hour notification of breakthroughs, emergence, or modification.

WORKED EXAMPLE — AIFM-BIO, SCIENTIFIC & BIOLOGICAL DESIGN MODELS

The reference implementation of the procedure

The highest-consequence cell in the matrix

AIFM-BIO sits closest to the Charter II prohibitions on AI-enabled biological and chemical weapon systems and on biological-threat optimization. The rubric measures governance-relevant outcomes only; it does not, and must not, enumerate harmful techniques. Assessment turns on five indicators, each measured at the elicitation ceiling.

B-1 · Dual-use marginal uplift
Meaningful assistance beyond the public frontier, reported as uplift against two baselines — not an absolute score.
B-2 · Design-capability generality
Whether the system is a bounded scientific instrument or a general biological-design capability.
B-3 · Safeguard robustness
Survival of refusal, input/output screening, and access controls at the ceiling — including after adversarial fine-tuning where weights could leave the developer’s control.
B-4 · Autonomy in the design loop
Any capacity to close a design-build-test cycle without a contemporaneous human authorisation gate.
B-5 · Access and reversibility
The deployment configuration under assessment — closed API, gated and screened access, open-weight, or on-device.

Determination follows the outcomes above: no measurable uplift and a gated loop → Certify with biosecurity screening and logging; bounded uplift with robust safeguards → Certify with conditions (licensed and screened access, KYC, six-month reassessment, open-weight excluded); material uplift or safeguard failure at the ceiling → Withhold; uplift crossing a Charter II threshold, or autonomous design directed at prohibited agents → Refer as Prohibited, with communication to competent biosecurity authorities.

Legal basis and publication. This procedure is set out in the IRBAI Evaluation Methodology, a subsidiary technical manual adopted and amended by the Executive Board through the Policy Forum process, subordinate in all respects to the Statute, Charter I, and the Annexes. The existence of the procedure, its indicators, and its determination consequences are published as Tier 1 (Public) material under Article 47. Evaluation battery contents, task-family specifics, and per-system measurements are withheld. — See also How We Regulate, the framework this procedure applies.