The AI Governance Stack
You’ve decided AI belongs in your GxP operations. The Stack is the architecture for that move. Traditional CSV assumes a fixed expected result, a system that stays put, behaviour independent of data, and visible failure; AI breaks all four. The Stack is how you keep control anyway.
Monitor & control
Drift, performance, human-in-the-loop and decision records.
Validation master plan
Risk-based and model-specific, not a one-size template.
Governance operating model
AI Governance Board, RACI, decision rights.
Classify & four-tier risk model
Tier the use case on impact and autonomy.
Reference architecture
Data, model, workflow, record, audit trail.
Regulatory spine
Annex 11 · draft Annex 22 · Part 11 · FDA CSA · FDA-EMA principles.
The four-tier GxP risk model
Classification is the most consequential hour in an AI system's lifecycle. The model tiers a use case by its impact on product quality and patient safety and by the system's autonomy, then sets the assurance accordingly. A worked example from the book: an HPLC chromatogram review classifier sits at Tier 2, and a case where a system drifted from Tier 2 to Tier 3 in eighteen months shows why reclassification triggers matter.
| Tier | Impact & autonomy | Assurance |
|---|---|---|
| Tier 1 | Low impact, assistive only | Light, with basic controls |
| Tier 2 | Significant GxP, human sign-off | Model validation and decision records |
| Tier 3 | High impact or higher autonomy | Full validation, tight monitoring, strong oversight |
| Tier 4 | Critical or autonomous in critical use | The heaviest controls, or kept out of that use |
Why AI breaks traditional CSV
| Traditional CSV assumes | AI reality |
|---|---|
| A fixed expected result | Results are probabilistic |
| A system that stays put | The model can change |
| Behaviour independent of data | Behaviour depends on training data |
| Failure is visible | Failure is often silent |
The layers, bottom up
Governance precedes validation. The regulatory spine (Annex 11, the draft Annex 22, 21 CFR Part 11, FDA CSA, the FDA-EMA principles) is the base. On it sits the reference architecture: data, model, workflow, record and audit trail. Then classification and risk tiering; then the governance operating model (an AI Governance Board, RACI and decision rights); then a risk-based, model-specific validation master plan; and at the top, monitoring and control for drift, performance, human-in-the-loop and decision-integrity records that capture input, output, reviewer and disposition.
Key takeaways
- Decision integrity is the new layer on top of data integrity.
- Classify first; the tier drives every downstream control.
- Monitoring is not optional: AI does not stay put, and failure can be silent.
Questions I get asked about the Stack
How do you validate an AI system in GxP?
Govern first, then validate. Classify the use case by its impact on product quality and patient safety and by its autonomy, then let the tier set the assurance: a model-specific validation plan, defined human oversight, and monitoring for drift once live. CSV alone cannot hold a system whose behaviour depends on data.
What does the draft EU GMP Annex 22 expect for AI in manufacturing?
The draft, published in July 2025, covers AI models in critical GxP applications: documented intended use, data quality, model validation with defined acceptance criteria, performance monitoring in operation, and human oversight, with audit trails recording which model version produced any given result. Each layer of the Stack maps to one of those expectations.
Why classify AI risk before validating?
Because the tier drives every downstream control. The four-tier model rates impact on product quality and patient safety against the system’s autonomy, and that rating decides how much validation, oversight and monitoring the use case carries. One deployed system drifted from Tier 2 to Tier 3 within eighteen months, which is why reclassification triggers matter.
