IRBAI · Certification Procedure
Certification is the procedure by which IRBAI assesses a model against the risk matrix and determines whether, and on what conditions, it may be certified for deployment. It is conducted under Chapter 14 of the Statute (Articles 100–109), against the mandatory assessment criteria of Annex I, and to the independent adversarial-testing standards of Annex II. The procedure produces a certification determination — a decision under Chapter 14. It does not set a system’s risk tier: the four tiers are fixed by Article 95 of the Statute.
This page describes the procedure. For the framework it applies — where regulatory risk sits — see How We Regulate.Capability is measured at the level a resourced and determined actor could elicit — not the level a system displays under default use. A measurement taken below that ceiling understates risk and is not accepted as evidence.
Independent evaluation by an Accredited Assessment Body takes precedence over a developer’s own claims. A material gap between submitted and independently measured capability is itself a certification finding.
Assessment yields one of four certification outcomes. These are decisions under Chapter 14; they do not alter a system’s statutory classification, and they are never averaged into a single number.
WHAT WE MEASURE — FOUR QUANTITIES, EACH EVIDENCING A STATUTORY CRITERION
Annex I fixes the mandatory criteria every system is assessed against. For foundation models, IRBAI gives those criteria evidential content through four measured quantities. Each is evidence toward a statutory criterion, not a substitute for it.
THE ELICITATION STANDARD
All capability and marginal-risk measurements are made at the level a resourced adversary could elicit. This operationalises Annex II Part 2, under which independent adversarial testing must cover, at minimum, prompt injection, data exfiltration, model exploitation, and “safety guardrail circumvention, including jailbreaking, refusal suppression, and capability elicitation beyond certified parameters.”
The assessment adversary is calibrated to the access mode — adversarial prompting; scaffolding and orchestration; tool and retrieval access; sampling at scale; and, wherever weights are or may become accessible, fine-tuning to remove or degrade safety behaviour. For any open-weight release, the ceiling always includes fine-tuning, because released weights place safety training within the recipient’s control.
THE EVIDENCE — IN FIXED ORDER OF PRECEDENCE
THE DETERMINATION — FOUR CERTIFICATION OUTCOMES
Assessment against a cell yields one of four outcomes. Where measurements span outcomes across the four quantities, the most restrictive outcome reached on any decision-relevant indicator governs; quantities are not averaged.
No material risk-relevant capability beyond the public frontier, or capability fully bounded by design.
Measurable risk-relevant capability, with mitigations demonstrated effective at the elicitation ceiling.
Material capability whose mitigations fail, or cannot be demonstrated, at the ceiling.
A capability falling within a Charter II category that cannot be reliably isolated or disabled (Article 95(b)).
RENEWAL AND REASSESSMENT
Certification is not permanent. High Risk (Tier C) certification is valid for no more than twelve months (Article 101); Systemic AI Systems additionally undergo independent adversarial testing at intervals not exceeding six months (Article 103). Reassessment is required earlier on any capability breakthrough, unexpected behavioural emergence, or material modification — with mandatory notification within twenty-four hours. A change in access mode, in particular any move toward open-weight release, and any containment breach are in themselves reassessment triggers.
WORKED EXAMPLE — AIFM-BIO, SCIENTIFIC & BIOLOGICAL DESIGN MODELS
AIFM-BIO sits closest to the Charter II prohibitions on AI-enabled biological and chemical weapon systems and on biological-threat optimization. The rubric measures governance-relevant outcomes only; it does not, and must not, enumerate harmful techniques. Assessment turns on five indicators, each measured at the elicitation ceiling.
Determination follows the outcomes above: no measurable uplift and a gated loop → Certify with biosecurity screening and logging; bounded uplift with robust safeguards → Certify with conditions (licensed and screened access, KYC, six-month reassessment, open-weight excluded); material uplift or safeguard failure at the ceiling → Withhold; uplift crossing a Charter II threshold, or autonomous design directed at prohibited agents → Refer as Prohibited, with communication to competent biosecurity authorities.