The first GMP-specific text on artificial intelligence, read for quality and manufacturing leaders.
Status as at 2 September 2026. Annex 22 remains in draft. Public consultation ran from 7 July to 7 October 2025. EMA's work plan targets Q4 2026 for providing final text to the European Commission. No publication date, implementation period or effective date has been announced.
The FDA and EMA joint principles set the direction of travel. Draft Annex 22 is where that direction becomes an inspection expectation for manufacturing.
It is short, roughly 6 pages, and it is the first GMP-specific text to address AI and machine learning models used in regulated production. It was drafted by the Inspectors Working Group of the EMA and PIC/S, and released as part of a package alongside a substantially revised Annex 11 on computerised systems and a revised Chapter 4 on documentation.
The gap between draft and final text is the opportunity because organisations that build governance and evidence now will not be retrofitting it under inspection pressure later.
Annex 22 applies to computerised systems used in the manufacture of medicinal products and active substances where AI models are used in critical applications with direct impact on patient safety, product quality or data integrity. It provides additional guidance to Annex 11 for systems with AI models embedded, and it applies to machine learning models that obtained their functionality through training with data rather than being explicitly programmed.
Two boundaries define the scope. The annex applies to static models, meaning models that do not adapt their performance during use by incorporating new data. Dynamic models that continuously and automatically learn and adapt during use are not covered by the document, and the draft states they should not be used in critical GMP applications.
The annex applies to models with a deterministic output, meaning identical inputs produce identical outputs. Models with probabilistic output are not covered, and again the draft states they should not be used in critical GMP applications.
Following from both, the draft states that it does not apply to generative AI and large language models, and that such models should not be used in critical GMP applications. Where they are used in non-critical applications, meaning applications without direct impact on patient safety, product quality or data integrity, the draft says personnel with adequate qualification and training should always be responsible for ensuring outputs are suitable for the intended use, in other words a human in the loop.
Why this matters. The wording is "should not" not "must not" and this is guidance rather than legislation. It is nonetheless the clearest signal available of what inspectors consider good practice. Many teams have adopted generative AI first, because it is visible and easy to trial. The draft pushes in the opposite direction. It is most permissive about the traditional machine learning that operations teams have run for years, and most restrictive about the generative tools attracting attention now. A governance framework that only addresses chatbot use will not survive a GMP inspection. One that ignores generative AI entirely will not reflect what staff are actually doing.
| Area | The expectation |
|---|---|
| Personnel | Close cooperation across process subject matter experts, QA, data scientists, IT and consultants, all with adequate qualifications, defined responsibilities and appropriate access. |
| Documentation | Documentation available and reviewed by the regulated user, whether the model was built in-house or supplied by a vendor. Supplier documentation supports the work. It does not transfer the responsibility. |
| Intended use | A detailed description of the task the model performs, a full characterisation of the input sample space including rare variations, and identification of limitations and possible erroneous or biased inputs. Approved by a named process subject matter expert before acceptance testing begins. |
| Acceptance criteria | Defined and approved before testing, with suitable test metrics, and set at least as high as the performance of the process the model replaces. That implies knowing how the existing process performs. |
| Test data | Representative, stratified, covering all subgroups, sufficient in size for statistical confidence, with labelling verified to a very high degree of correctness. Generating test data or labels using generative AI is not recommended and any use must be fully justified. |
| Test data independence | Test data not used in development, training or validation, protected by access control and audit trail, with no copies outside the repository. Staff who have had access to test data should not be involved in training the same model, or should work in pairs under a 4-eyes principle where that is impossible. |
| Test execution | An approved test plan before testing, evidence the model generalises well rather than overfitting, and documented investigation of any deviation or failure to meet acceptance criteria. |
| Explainability | For models in critical applications, capture and record which features in the test data drove a particular classification or decision, using techniques such as feature attribution or heat maps, with feature review built into approval of test results. |
| Confidence | Log the confidence score for each prediction or classification where applicable, set an appropriate threshold, and consider flagging an outcome as undecided rather than returning an unreliable one. |
| Operation | Change control and configuration control before deployment, regular monitoring of model performance and of drift in the input data, and records kept of human review where a person is making the decision the model informs. |
It does not stand alone, and reading it as a complete AI compliance framework is the most common mistake we see.
Annex 22 was released as one of three linked drafts, alongside a revised Annex 11 on computerised systems and a revised Chapter 4 on documentation. It provides additional guidance to Annex 11 rather than replacing it. Everything an AI model does not cover, from data integrity and record retention through to eventual system retirement, sits in those documents and in EU GMP more broadly. If your AI governance work stops at Annex 22, you will have documented your models and left the system around them untouched.
It is also not the only applicable framework. The EU AI Act applies horizontally to providers and deployers of AI systems in the EU. A critical application under Annex 22 is not automatically a high-risk AI system under the AI Act. The two classifications, and the obligations that follow from each, need to be assessed separately.
On 2 April 2026 the FDA issued a warning letter to a Michigan manufacturer of homeopathic drug products containing a section headed "Inappropriate Use of Artificial Intelligence in Pharmaceutical Manufacturing" the first time AI has been cited as a standalone deficiency. The firm had used AI agents to create specifications, procedures and master production and control records. Asked why process validation had not been performed before distribution, it said the AI agent had never raised the requirement.
Why this matters.Annex 22 puts documented intended use, independent test data and human accountability onto the inspector's list in Europe. The FDA has already made a finding on the same principle in the United States without waiting for AI-specific guidance. This approach does not look likely to soften.
The ISPE GAMP Guide: Artificial Intelligence (2025) is the practitioner companion. Annex 22 states the regulatory expectation. The GAMP guide describes how to meet it through risk-based validation, supplier assessment and lifecycle controls that extend the familiar GAMP 5 approach. Teams already working to GAMP 5 usually find the step smaller than expected. The gap is typically AI literacy and governance ownership, not validation capability.
The Institute of Applied AI helps life sciences organisations build AI capability, assign clear governance to AI decisions, and turn ambition into roadmaps that survive inspection. If you would like to discuss any of the above, get in touch.
Sources: Draft EU GMP Annex 22: Artificial Intelligence, consultation text (7 July to 7 October 2025) , EMA GMP Inspectors Working Group work plan 2026, ISPE GAMP Guide: Artificial Intelligence (2025), FDA and EMA Guiding Principles of Good AI Practice in Drug Development (January 2026), FDA Warning Letter, Purolea Cosmetics Lab, 2 April 2026
Annex 22 is draft guidance and its final wording may change. Nothing here is a compliance position or legal advice.