GAMP5 Risk Classification for AI/ML Models: Are You Applying the Right Validation Depth?
GAMP5 wasn't built for adaptive AI/ML algorithms. Here's how to assign risk classification correctly and build revalidation triggers that hold up to FDA scrutiny.
Most pharmaceutical labs have a rehearsed answer when an FDA investigator asks how they validated their AI-powered LIMS module: “We treated it as a Category 5 custom system and ran full IQ/OQ/PQ.” The investigator nods — then asks for the revalidation protocol that followed the last model weight update. The silence that follows is expensive.
This is the real gap in how regulated industries are handling AI/ML software right now. Not the initial validation. The ongoing governance. GAMP5’s second edition, published by ISPE in 2022, made meaningful progress — but the framework was designed for deterministic software. Adaptive algorithms break several of its foundational assumptions, and regulators are starting to notice.
Why GAMP5’s Category System Wasn’t Built for Adaptive Algorithms
GAMP5 organizes software into four active categories based on complexity and customization. Category 1 covers infrastructure software — operating systems, network tools, virtualization layers. Category 3 covers non-configured commercial off-the-shelf (COTS) software used as delivered, without site-specific modification. Category 4 covers configured COTS systems: think an ERP or LIMS with site-specific workflow rules, user access tiers, and business logic applied through configuration. Category 5 is reserved for custom software, where your organization (or a contractor working on your behalf) authored the application code.
Traditional validation fits this structure cleanly. A configured LIMS is Category 4. You document what you configured, qualify the functions you’ve enabled, and rely on supplier qualification for the underlying vendor-developed code. A homegrown batch release calculation script? Category 5 — full IQ/OQ/PQ, source code review, the works.
AI/ML models don’t fit neatly here. A machine learning model that classifies out-of-specification results, flags anomalies in chromatographic data, or predicts equipment failure might be delivered as a vendor module — technically Category 4 in terms of its provenance. But its functional behavior is determined by training data and model weights, not by configuration settings. Change the training dataset and you’ve effectively changed how the software behaves, even if no lines of code were touched and the vendor never issued a patch.
GAMP5’s 2022 edition acknowledged this tension and introduced supplementary AI/ML considerations, directing practitioners toward a risk-based approach within the existing categories with enhanced change control for adaptive components. That’s technically sound guidance. It’s also thin enough to leave validation teams rewriting protocols the night before an inspection.
The Category 4 vs. Category 5 Question — And Why It Matters More Than You Think
Category assignment isn’t bureaucratic box-checking. It determines the entire validation burden: testing depth, documentation volume, required supplier involvement, and — critically — what constitutes a “change” that triggers revalidation. Getting this wrong leaves your organization exposed to 483 observations that are slow and costly to remediate.
Here’s where many regulated companies make an expensive classification error. Teams assign a vendor-supplied AI/ML module to Category 4 because the vendor wrote the underlying model architecture. That seems reasonable on its face. But if your organization contributes the training data, fine-tunes the model on your historical batch records, or defines the decision thresholds — that’s your accountability, not the vendor’s. The functional risk profile looks much more like Category 5, regardless of who holds the source code intellectual property.
FDA’s draft Computer Software Assurance (CSA) guidance, published in September 2022, reframes this productively. The CSA framework — which FDA intends to eventually supersede traditional CSV — shifts emphasis from protocol execution to critical thinking about risk and testing evidence. Under CSA logic, the right question isn’t “what category applies?” It’s “what is the patient safety and data integrity risk if this AI component fails, and how would we detect it before that failure reaches a batch record or a released product?”
That reframing changes the validation approach materially. For an AI model making decision-grade outputs — flagging a batch for rejection, recommending a stability test interval, or classifying an audit finding — the risk-based answer almost always demands Category 5 rigor, regardless of how the software arrived. Specifically: documented performance benchmarks established at validation against a held-out test dataset, defined acceptable performance boundaries with explicit tolerances, and predefined revalidation triggers when performance drifts outside those boundaries.
The number most validation packages are missing: a quantitative baseline metric — F1 score, accuracy, recall, or whatever measure is appropriate to your application — with a documented ±15% tolerance band. If your model’s false-positive rate on OOS detection drifts by more than that threshold on a rolling 30-day basis, you need a revalidation event. Not a deviation log entry. A revalidation event.
How FDA’s CSA Framework Changes Your Documentation Strategy
The September 2022 CSA draft guidance emerged from years of industry frustration that traditional CSV was generating enormous documentation overhead without proportionate quality benefit. FDA observed — correctly — that flawlessly executed test scripts don’t prove a system works. They prove a script was executed. Actual test results prove a system works. The distinction sounds subtle. Its practical implications are significant.
For AI/ML systems, this matters more than anywhere else in GxP software. An IQ/OQ/PQ protocol that verifies installation completion and confirms model training converged is largely theater from a quality assurance standpoint. The meaningful validation question is whether the model performs correctly on data representing your actual operating conditions — including edge cases, seasonal variation in incoming raw material quality, and operator handling patterns that differ from your curated training set.
Under CSA principles, three documentation elements have moved from nice-to-have to audit-critical for any AI component in a GxP-regulated workflow:
Intended Use Statement. Not the vendor’s product description. Your organization’s specific, bounded statement of what the AI does and doesn’t do in your process — including explicit scope limitations. “This model classifies tablet dissolution results against specification limits for Product A, Lots 1000–9999, using USP Apparatus II methodology at 37°C” is defensible. “AI-assisted QC tool” is not.
Training Data Governance Record. Where did the training data originate? How was it qualified for training use? Who reviewed and approved it? Data integrity requirements under 21 CFR Part 11 apply to training datasets just as they apply to batch records or stability study raw data. Audit trail, review records, version control — the same standards. This surprises many validation teams. It shouldn’t.
Model-Specific Change Control Procedure. Your standard software change control SOP was designed for versioned product releases. A model retraining on six additional months of production data isn’t a software release — but it changes functional behavior. Your change control procedure needs criteria that explicitly distinguish a model recalibration within validated bounds (minor change) from a model retraining on new data or an architecture modification (major change requiring revalidation). Most quality systems have the generic procedure. Almost none have the AI-specific criteria written into it.
Building a Revalidation Trigger Framework That Holds Up to Scrutiny
Static software has a validated state — a specific version, configuration, and installation that matches what was qualified. AI/ML systems have a validated performance envelope — a defined range of acceptable behavior that the system must remain within across its operating lifetime. These two things require fundamentally different ongoing governance approaches, and conflating them is where most quality systems currently fall short.
A practical revalidation trigger framework for a GxP AI/ML system should address four distinct trigger categories:
Performance drift triggers. Predefined statistical thresholds against your quantitative baseline. When the primary performance metric drops below the lower control limit established at validation — typically ±15% for most pharmaceutical decision-support applications, tightened for safety-critical outputs — a revalidation event is triggered automatically. Not flagged for managerial discretion. Triggered.
Input data distribution triggers. If the data the model processes in production shifts meaningfully from the distribution it saw during training — measurable through statistical monitoring on input feature characteristics — the model may be operating outside its validated domain even before performance metrics visibly degrade. This is particularly common when manufacturing processes change, new raw material suppliers are qualified, or an equipment upgrade alters sensor calibration ranges.
Model update triggers. Any change to model weights, training data composition, architecture, feature engineering, or inference thresholds flows through change control with a documented impact assessment. The assessment classifies the change and determines whether re-qualification testing or full revalidation is required. Document the classification criteria explicitly. “Significant changes will be evaluated” will earn a 483 observation. “Changes resulting in greater than 10% shift in primary performance metric require full revalidation” will not.
Regulatory and intended use triggers. If FDA issues new guidance relevant to your AI application — and FDA’s AI/ML Action Plan, first published in January 2021 and updated since, signals that more guidance is coming — or if your organization expands what the system is used to decide, the existing validation scope may be insufficient regardless of where performance metrics sit.
Build all four trigger categories into your Computerized System Validation Plan — or, if you’re transitioning to CSA, into your Software Assurance Plan. Attach measurable, objective criteria to each. Attach responsibility assignments. And run a tabletop exercise once a year to confirm the trigger criteria still match your actual operating conditions.
What to Do Before Your Next Inspection Cycle
Pull your current validation documentation for every AI/ML-assisted system in a GxP-regulated workflow. Check for three specific elements.
First: a quantitative performance baseline, documented at initial validation, expressed in a unit that a model monitoring system can track automatically on a continuous or periodic basis. Not a narrative description of expected behavior. A number, with a tolerance band, tied to a defined dataset.
Second: a training data governance record that satisfies 21 CFR Part 11 requirements — audit trail, review and approval records, version control, and a qualification rationale for why that dataset was representative of your operating conditions at the time of validation.
Third: a change control procedure that distinguishes between model recalibration, model retraining, and model architecture changes, with defined criteria for each classification and the corresponding revalidation scope.
If any of the three are missing, you have an active gap. GAMP5’s second edition gives you useful scaffolding; FDA’s CSA guidance tells you what the evidence needs to actually prove. For AI/ML systems, those two frameworks together still require your organization to supply the model governance layer — because no published standard can define what acceptable AI performance looks like for your specific product, process, and risk tolerance. That definition is your responsibility. And it’s the first thing an investigator will ask for.
Written by Sam Sammane, Founder & CEO, Aurora TIC | Founder, Qalitex Group. Learn more about our team
Reserve early access to our AI audit tools Contact us
Related from our network
- ISO 17025-accredited laboratory testing for pharmaceutical and nutraceutical manufacturers — Qalitex Laboratories provides analytical testing and compliance support for US-regulated facilities, including LIMS-integrated COA generation.
- Health Canada GMP compliance and NHP testing services — Androxa supports Canadian manufacturers navigating NHPD and pharmaceutical GMP requirements, with validated testing workflows aligned to Health Canada expectations.