What operating model should pharma quality use to govern AI?
Define the context of use before evaluating the model, assess risk based on decision influence and error consequence, apply documented data and model controls, require active human oversight, set performance acceptance criteria against a known baseline, and monitor for drift after deployment — adapted from FDA's January 2025 risk-based credibility assessment framework.
Most AI governance frameworks start with the model — what it is, how accurate it is, what it can do. FDA's own framework starts somewhere else entirely: what decision is this actually informing, and who is accountable for it?
Context of Use
Before any technical evaluation, define the model's context of use: the specific question it answers, who acts on its output, and whether it makes an autonomous decision or informs a human one. This is the second step in FDA's January 2025 credibility assessment framework, and it comes before risk assessment for a reason — the same model architecture can be low-risk in one context and GxP-critical in another. A classification model flagging documents for human review carries different risk than the same model auto-routing a batch disposition. Evaluating the model before defining what it's actually used for gets the sequence backwards.
Risk and Credibility Assessment
FDA frames model risk as a combination of two factors: how much the model's output influences a decision, and how severe the consequence would be if that output were wrong. A model providing one input among several factors a human weighs carries lower risk than one whose output is acted on directly with no further check — even when both models perform an identical underlying calculation. Credibility assessment then defines what evidence is needed to trust the model for that specific context, scaled to the risk level established here, not a fixed checklist applied uniformly to every model regardless of stakes.
Data and Model Controls
Training data provenance, representativeness, and version control need to be documented well enough that someone can explain what the model actually learned — because if you can't explain that, you can't defend what it produced when questioned. This includes tracking which model version is in production, what data trained it, and maintaining that record through any retraining event, not just at initial deployment.
Human Oversight
Oversight means a qualified person actively reviews and takes documented responsibility for AI output before it's used, not a passive approval step that exists on paper. FDA's first AI-specific warning letter, issued in 2026, cited a manufacturer for using AI-generated specifications and procedures without this kind of review — the same failure this governance step exists to prevent. Active oversight means checking output against source truth and documenting what was verified, the same standard applied to any other GxP record.
Performance Acceptance Criteria
You need the baseline before you can set the bar. An AI model's acceptance criteria should be set against the known performance of whatever it's replacing or supporting — a manual process, an existing method, a human reviewer's typical accuracy. Many organizations haven't formally measured that baseline, which makes it impossible to state with any rigor that the AI model actually performs adequately for its context of use, rather than just performing plausibly.
Drift Monitoring and Change Control
A model's performance at validation reflects the data available at that point in time. Real-world data can shift afterward in ways that degrade accuracy without any visible change to the system itself — the same structural problem continued process verification solves for a manufacturing process, applied to a statistical model instead of a process parameter. Drift monitoring tracks performance against the original acceptance criteria on an ongoing basis, and any retraining or model update needs to go through change control just as a configuration change would, with the same documented impact assessment.
Required Validation Evidence
| Evidence | What It Confirms |
|---|---|
| Context of use statement | What decision the model informs, and who's accountable for acting on it |
| Risk assessment | Decision influence and error consequence, driving required evidence depth |
| Credibility assessment plan | What evidence will be gathered to establish trust for this specific use |
| Execution results | The plan actually carried out, with results documented as they occurred |
| Deviations from plan | Any departure from the credibility plan, explained rather than omitted |
| Adequacy determination | An explicit conclusion on fitness for the defined context of use |
How GoVal Supports AI Governance for Pharma Quality
GoVal manages this operating model as a structured, risk-assessed record — tracking each AI system's documented context of use, risk classification, human oversight sign-off, and performance acceptance criteria in the same audit-trailed structure used for conventional GxP systems. Drift monitoring triggers are tied to change control, so a performance degradation flags reassessment automatically rather than depending on someone noticing during a periodic report review.
Related Topics
Frequently Asked Questions
What is AI governance for pharma quality, and what does it require? +
What is FDA's AI credibility assessment framework? +
What is "context of use" in AI governance, and why does it come first? +
How is AI model risk assessed differently from software risk generally? +
Why does AI governance require ongoing drift monitoring after deployment? +
Does FDA's AI credibility framework apply to AI used for internal drafting or operational tasks? +
How does GoVal support AI governance for pharma quality? +
Govern AI-embedded systems with the same rigor as everything else
Context of use, risk classification, oversight sign-off, and drift-triggered change control — in GoVal.
