What an inspector asks about your AI
The question is never "is your model accurate". I have taken AI tools in live GxP quality operations through health authority inspections, and accuracy came up as a subheading at most. What inspectors want is something harder to fake: show me how you decided this much evidence was enough, and show me it is still true today.
Six things sit underneath that. Each has to be shown rather than described, and a team that can produce all six on demand will not be caught out by whatever the final regulatory text says.
A word on where the rules stand. Draft Annex 22 on artificial intelligence is in consultation, and I contribute to the industry review of it, which mostly means I have read the arguments more times than is healthy. The wording may still move. None of the six below depends on that, because all six are already good validation practice under Annex 11, 21 CFR Part 11 and GAMP 5.
1. Intended use, and the boundary
What is this model allowed to decide, and what must it never decide? Write that as two sentences a non-specialist can read.
This is the most common gap I find, and the most expensive one, because everything else hangs from it. The risk tier follows from intended use. The evidence scope is sized by it. The monitoring plan exists to protect it. When the statement is missing, or when it lives in a vendor's documentation rather than your own quality system, every downstream control rests on nothing.
A thin answer sounds like: "It helps with deviation handling." A good one sounds like: "It proposes a preliminary classification for deviations in categories A and C, which a trained investigator must confirm or override before the record advances. It does not close deviations, and it is not used for categories B or D."
2. Data lineage
Where did the training data come from, and why does it represent the process you are applying the model to?
Inspectors follow data. Expect questions about the source systems, the period covered, how the data was cleaned and by whom, and whether anything was excluded. The one that catches teams out is representativeness: a model trained on two years of data from three sites, then deployed at a fourth site with different equipment and a different batch profile, is operating outside the evidence you hold, whether or not anyone noticed.
3. Performance evidence
What acceptance criteria did you set, when did you set them, and did the model meet them on data it had never seen?
The order matters. Criteria set after seeing the results are not criteria, and an experienced reviewer will spot it in the dates. Hold back a genuine test set, include the difficult and rare cases rather than only the clean ones, and record the failures alongside the successes. A validation pack showing 100% success on comfortable data is less convincing than one showing where the model struggles and what you decided to do about it.
4. Human oversight
Who reviews the output, what competence qualifies them to do it, and what are they permitted to overrule?
Human oversight fails quietly in two ways. Either the reviewer has no realistic way to judge the output, so the review becomes a rubber stamp, or the workflow makes overriding the model slower than accepting it, so nobody ever does. Both show up in the data if an inspector looks, and increasingly they do look. If your override rate is zero across thousands of transactions, that is not a sign the model is perfect.
Train the reviewers, give them a route to disagree that costs them nothing, and monitor how often they use it.
5. Change control
Every version, every retraining event, every threshold adjustment, traceable to who approved it and why.
This is where AI departs from conventional software, and where good validation teams still get caught. A retrained model is a changed system even when the code is untouched and the vendor calls it an update. A threshold moved from 0.8 to 0.75 changes behaviour without changing a single line. If your change control cannot see these, it cannot govern the thing you are actually running.
6. Ongoing monitoring
How would you know the model had drifted, and what happens on the day it has?
A validated model is valid for the data it met. Then a supplier changes a form, a site joins, a season turns, and performance slips inside a tolerance nobody is measuring. The tell is rarely a metric. It is users quietly working around the output, and that never appears in a report.
What holds up under questioning is modest and specific: performance watched monthly against the criteria you set at release, retraining and threshold changes routed through change control, and an annual question that catches more than the monitoring does, which is whether the intended use is still the actual use.
The honest summary
If you cannot produce these six on demand today, that gap is your work for this quarter, and none of it requires the draft to be final. If you can, you are ahead of most of the organisations I walk into, and the conversation with an inspector becomes a professional one rather than an anxious one.
Questions leaders ask
What will an inspector ask about AI in a GxP process?
Six things, and each has to be shown rather than described: the intended use and its boundary, the lineage of the training data, performance evidence against criteria set before testing, who provides human oversight and what they may overrule, change control covering every version and retraining event, and how drift is detected and acted on.
Is EU GMP Annex 22 final?
No. Draft Annex 22 on artificial intelligence is in consultation and the final wording may change. The six expectations discussed here are worth building for now because they are already good validation practice and do not depend on the final text.
Do we need to validate AI that only drafts documents?
It depends on what the output becomes. If a human reviews and owns the result before it enters a GxP record, the system is qualified for its intended use with the review documented. If the output enters a record without meaningful review, it is making a GxP contribution and owes a great deal more evidence.
What is the most common gap in AI validation evidence?
The intended-use statement. It is often missing, or it lives in a vendor document rather than in the company's own quality system. Every other control hangs from it, so when it is absent the risk tier, the evidence scope and the monitoring plan all rest on nothing.
The method, with the templates and the question bank
Validating AI in GxP works all six of these through in full, with the evidence patterns per risk tier and the inspection questions to expect. Rated 4.8 out of 5 by its first readers.
Get the book Train your teamNot sure where you stand? The AI Readiness Check takes six minutes and names the first gap to close.
