Annex 22 is still a draft. Here's what you can already do with AI in GMP


If you work in quality, production or validation, you have probably heard some version of this: "Europe is going to ban AI in manufacturing." It usually comes up in conversations about the draft EU GMP Annex 22, which went out for public consultation between July and October 2025. Read the text and a different picture emerges. The draft does not ban AI. It draws a fairly clear line between where a model can make the call and where it can only help a person make it. That line is useful today, even with the final text still open.
Where Annex 22 stands today
- Annex 22 is a draft. The European Commission opened the stakeholder consultation on 7 July 2025 and closed it on 7 October 2025, together with draft revisions of Chapter 4 (Documentation) and Annex 11 (Computerised Systems). All three documents were drafted by the EMA GMDP Inspectors Working Group in cooperation with PIC/S.
- There is a target, not an effective date. The working group's 2026–2028 work plan aims to provide the Commission with a final Annex 22 text in Q4 2026. No adoption or effective date has been announced.
- PIC/S has not adopted Annex 22. The current PIC/S GMP Guide is PE 009-18.
In short, the text may change. Everything below refers to the draft consulted on in July–October 2025.
Decide or assist: what the draft actually says
The starting point is scope. The draft applies to AI models used in critical applications, meaning those with a direct impact on patient safety, product quality or data integrity, for example to predict or classify data. It then makes three distinctions:
- Static models, which do not change their behaviour during use, are in scope. Dynamic models, which keep learning in operation, are not, and "should not be used in critical GMP applications".
- Deterministic output (same input, same output) is in scope. Models with probabilistic output, which may answer differently to the same input, should not be used in critical applications either.
- Generative AI and LLMs are out of scope and should not be used in critical GMP applications. In non-critical applications, adequately qualified and trained personnel should always be responsible for ensuring the output is suitable for its intended use: human-in-the-loop.
Note that the draft uses "should" throughout, including in the "should not" statements above.
For models used in critical applications, the draft asks for evidence, for example:
- intended use and input sample space described in detail, including common and rare variations, owned by a process subject matter expert (SME) and approved before acceptance testing starts;
- acceptance criteria at least as high as the performance of the process the model replaces, which means you need to know that performance;
- test data that is representative and independent, never used in development, training or validation, and protected by access control and audit trail. Generating test data or labels (for example with generative AI) is "not recommended" and, if done, needs full justification;
- explainability: capturing, during testing, which features drove a given classification (for example through feature attribution);
- confidence: logging, during testing, the confidence score for each prediction; in operation, where the score is very low, considering flagging the result as "undecided";
- operation: change control, configuration control, performance monitoring, input-data drift monitoring, and records of human review where the model feeds an operator's decision.
The draft Chapter 4 is also clear on accountability. Under 4.24, accountability for the integrity of records produced or processed with AI rests with the regulated user. Under 4.25, that kind of automated support should be included in the pharmaceutical quality system, whether the service runs on premise or is hosted.
Put together: where the impact is direct, the model has to be predictable and proven. Where generative AI comes in, it comes in as an assistant, with a qualified person answering for the result.
The open question: generative AI
This is where the final text may differ from the draft. According to EMA, the 2025 consultation suggested support for potentially enabling technologies such as generative AI and LLMs in medicines manufacturing. The agency says it is still considering the implications. On 30 June and 1 July 2026 it held an expert workshop on control and mitigation measures such as guardrails, as part of a proposed risk-based approach. One of the questions put to experts was whether there are classes of critical decisions where no combination of guardrails and oversight would be enough. The workshop report came out on 2 October 2026. Participants argued for a risk-based, technology-neutral approach, and the drafting group will now assess that input as it revises Annex 22. In other words, the final text is still open.
In the US, FDA's January 2025 guidance on using AI to support regulatory decision-making for drugs and biological products is still a draft. In 2023, CDER published a discussion paper on AI in drug manufacturing, which is explicitly not guidance.
Some final documents do help: the ISPE GAMP Guide: Artificial Intelligence, published in July 2025, and the Guiding Principles of Good AI Practice in Drug Development, published by EMA and FDA on 14 January 2026. These are ten high-level principles rather than requirements, and they include a risk-based approach, a clearly defined context of use, and life cycle management with drift monitoring.
Whatever EMA decides on generative AI, the draft, the EMA–FDA principles and the workshop questions all point the same way: know where AI is being used, at what level of risk, and who is accountable for it.
What you can do now, without waiting for the final text
What the draft asks of models that decide and of models that assist shares a common foundation, and that foundation is likely to survive revisions. In practice, it comes down to four steps.
Take inventory. There is often more AI in use than people think: a vision model on packaging inspection, predictive maintenance on equipment, an assistant that summarises deviations, a suggestion feature inside a third-party application. List them all and ask the scope question for each: is there a direct impact on patient safety, product quality or data integrity?
Write the intended use before you test. For a critical use, that means input space, known limitations and acceptance criteria, with SME approval. For a non-critical one, a short description of what the tool is for and who is responsible for its output already helps.
Make the accountable person explicit. Where a model feeds an operator's decision and testing of the model has been reduced, the draft expects the operator's responsibility to be part of the intended use, their training and performance to be monitored like any other manual process, and the review to be recorded.
Bring suppliers into scope. Under the draft, model documentation should be available to and reviewed by the regulated user, whether the model was trained in-house or by a supplier. That is worth checking before you sign the contract.
Checklist: five questions for every AI use
The four steps above fit into five questions you can run item by item through the inventory:
- Critical or non-critical? Is there a direct impact on patient safety, product quality or data integrity? If so, the draft covers static, deterministic models backed by evidence, and says dynamic, probabilistic and generative models should not be used there.
- Are intended use and input space written down? With limitations identified and SME approval before acceptance testing.
- Is the test data independent? Kept apart from development, training and validation, representative, under access control and audit trail. Generating test data or labels by any means (generative AI being one example) is not recommended; if done, it needs full justification.
- Is human review recorded? Who reviews, with what qualification, and where each approval is recorded.
- Are change control and drift monitoring in place? The model is under change and configuration control, with performance and input-data drift monitored.
How T2 approaches AI
Those five questions are also the criteria we use ourselves. At T2, we apply AI according to risk. In our view, generative AI is there to assist, with a qualified person approving. Where the decision is critical, the logic is static, deterministic models, backed by evidence and under change control. That is the criterion we use to assess any use of AI in regulated processes such as electronic batch records, weighing and dispensing, and OEE.
Draft Annex 22 may still change. Keeping the right person accountable for the right decision, we believe, will not.
References
- European Commission – Stakeholders' Consultation on EudraLex Volume 4: Chapter 4, Annex 11 and New Annex 22 (7 July – 7 October 2025)
- Draft Annex 22 – Artificial Intelligence (consultation document)
- Draft revised Chapter 4 – Documentation
- Draft revised Annex 11 – Computerised Systems
- EMA – 3-year work plan for the GMDP Inspectors Working Group (January 2026 – December 2028)
- EMA – GMDP Inspectors Working Group (work plan and 2025 annual report)
- EMA – Multistakeholder workshop on expert contributions to AI guidance development (Annex 22), 30 June – 1 July 2026
- EMA – Report of the multistakeholder workshop on expert contributions to AI guidance development (Annex 22), published 2 October 2026
- PIC/S – Publications (PIC/S GMP Guide PE 009-18)
- FDA – Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products (draft guidance, January 2025)
- FDA/CDER – Artificial Intelligence in Drug Manufacturing (discussion paper, 2023)
- EMA–FDA – Guiding Principles of Good AI Practice in Drug Development (January 2026; EMA-hosted PDF)
- EMA – EMA and FDA set common principles for AI in medicine development (14 January 2026)
- ISPE – GAMP Guide: Artificial Intelligence (July 2025)
Practical writing on GxP, MES, data integrity and shop-floor systems. A few times a month, no noise.


