What is EU GMP Annex 22 and what does it require?
Annex 22 is the EU's first dedicated framework for artificial intelligence in GMP manufacturing, published as a draft on July 7, 2025. It applies to static, deterministic AI/ML models embedded in critical applications affecting patient safety, product quality, or data integrity, requiring documented intended use, acceptance criteria no lower than the process replaced, independent test data, and change control tied to retesting decisions. Generative AI and continuously adapting models are excluded from critical use.
Annex 22 doesn't ban AI from GMP manufacturing. It bans the kind most teams are quietly piloting right now — and the fine print on staff independence for test data is stricter than almost anyone expects.
What Was Published, and When
On July 7, 2025, the European Commission published a draft of Annex 22, a brand-new addition to EudraLex Volume 4, alongside a substantially revised Annex 11 and an updated Chapter 4 on documentation. Stakeholder consultation closed in October 2025, and final adoption is expected sometime in 2026. Annex 22 doesn't replace Annex 11 — it supplements it, applying specifically to computerized systems in which AI or machine learning models are embedded and used in ways that could directly affect patient safety, product quality, or data integrity.
Scope: Which AI Actually Falls Under Annex 22
The scope is narrower than most "AI governance" conversations assume, and getting this distinction right is the first real decision point.
| Model Type | Status Under the Draft |
|---|---|
| Static, deterministic models (trained and frozen) | Covered — permitted in critical applications if validated per Annex 22 |
| Dynamic models (continuously learn during use) | Excluded — not covered, and should not be used in critical GMP applications |
| Probabilistic models (same input, variable output) | Excluded — not covered, and should not be used in critical GMP applications |
| Generative AI and large language models | Excluded from critical use — permitted only in non-critical, human-in-the-loop contexts |
This rules out most of what people currently mean by "AI" in casual conversation. It's aimed squarely at classification and prediction models — the kind already piloted in manufacturing today, such as neural-network models that predict tablet potency in real time from process data instead of waiting on lab testing, computer vision systems classifying visual defects on a fill line, or predictive maintenance models flagging equipment likely to fail before it affects a batch. Each of these is a static, deterministic classifier making a GMP-critical call — exactly Annex 22's target, and exactly the kind of system that needs a documented intended use, acceptance criteria, and independent test data before deployment.
The Requirement Most Teams Will Miss: Test Data Independence
Annex 22 requires genuine separation between the data used to train and validate a model and the data used to test it — not just a different file, but, where feasible, different people. Staff who developed or trained a model shouldn't also be the ones testing it against the held-out dataset; where that separation isn't organizationally possible, a four-eyes approach is required, pairing someone with prior access to the test data with a colleague who has none. Access to test data itself needs audit trail protection, logging who touched it and when.
Why this matters more than it sounds: most pharma quality teams have never had to formally separate "who built it" from "who tested it" at the level of individual staff access to a specific dataset. This is closer to a clinical-trial-style blinding control than a typical CSV requirement, and it has real staffing and process implications for smaller teams where the same two or three people currently do all of it.
The "No Decrease" Principle
An AI model's acceptance criteria must be set at least as high as the performance of the process it's replacing — which means that process's performance has to already be known and quantified. This is a genuinely common gap: many manual inspection or classification steps a company wants to automate have never had their own accuracy, sensitivity, or error rate formally measured. Under Annex 22, that baseline isn't optional context — it's the number the model has to beat before deployment is justified at all.
If You're Piloting GenAI Copilots Right Now
A lot of pharma teams currently experimenting with LLM-based copilots — drafting deviation reports, summarizing batch records, searching SOPs for a relevant procedure, drafting maintenance notes — should read the scope section carefully. Annex 22 is explicit that generative AI and LLMs don't qualify for critical GMP use under its framework, because they don't produce deterministic output. That doesn't prohibit their use outright: drafting a first-pass deviation summary or pulling up a relevant SOP section stays firmly non-critical, provided a qualified person reviews the output and takes documented responsibility for any action taken. What crosses the line is using that same LLM to decide whether a batch passes or fails, or to independently classify an out-of-specification result — that's a GMP-critical decision, and Annex 22 says a non-deterministic model can't be trusted with it, full stop.
What Annex 22 Does Not Change
It doesn't shift validation responsibility to the vendor. Whether a model is built in-house or supplied by a third party, the documentation for intended use, acceptance criteria, test data, and testing must be available and reviewed by the regulated company itself. A vendor's own internal validation work is useful supporting evidence — the same way a SOC 2 certificate supports vendor qualification for a SaaS platform — but it doesn't transfer accountability for whether the model is fit for your specific GMP-critical use.
Regulatory Timeline
Draft published July 7, 2025. Stakeholder consultation closed October 2025. Final adoption is expected during 2026, with enforcement likely phased in over the following one to two years based on how comparable EudraLex revisions have historically rolled out. The exact publication date and grace period aren't confirmed as of this writing, so treat the draft's substance as the working baseline for preparation now, not something to wait on.
How GoVal Supports Annex 22 Readiness
GoVal extends the same risk assessment, change control, and audit trail structure it already applies to conventional GxP systems to cover AI model validation evidence — intended use documentation, acceptance criteria, and test data access records live in one auditable structure rather than scattered across data science tooling and a separate document repository. When a model, its host system, or the process it automates changes, GoVal's change control workflow flags whether retesting is required and captures the justification either way, generating exactly the kind of defensible record Annex 22 is built around.
Related Topics
Frequently Asked Questions
What is EU GMP Annex 22? +
How does Annex 22 relate to Annex 11? +
When will Annex 22 be finalized? +
Does Annex 22 apply to ChatGPT or other large language models used in pharma? +
What AI models are prohibited from critical GMP applications under Annex 22? +
Does Annex 22 apply outside the EU? +
How does GoVal support Annex 22 compliance for AI-embedded systems? +
Get ahead of Annex 22 before it's finalized
Risk assessment, change control, and audit-trailed evidence — extended to your AI-embedded GxP systems, in GoVal.
