What should an AI risk assessment for GxP systems include?
Context of use, model type, decision influence, consequence of error, existing compensating controls, required assurance activity, human oversight requirement, baseline performance reference, and a reassessment trigger. The risk rating comes from combining decision influence with error consequence — not from how sophisticated or accurate the model is in isolation.
Two AI models can use the exact same underlying algorithm and land at opposite ends of the risk scale. The model isn't what you're rating — the decision it feeds, and what happens if that decision is wrong, is.
The AI Risk Assessment Template
| Field | What It Captures |
|---|---|
| Context of use | What question the model answers, who acts on the output, autonomous or advisory |
| Model type | Static/deterministic, dynamic/adaptive, or generative |
| Decision influence | Sole basis for action, or one input among several a human weighs |
| Consequence of error | Impact on patient safety, product quality, or data integrity if the output is wrong |
| Compensating controls | Independent human review, secondary checks, or other safeguards already in place |
| Risk rating | Combination of decision influence and consequence — the output of this assessment |
| Required assurance activity | Scripted testing, scenario-based testing, or vendor evidence, scaled to the rating |
| Human oversight requirement | Mandatory review before use, or autonomous operation permitted |
| Baseline performance reference | What the model is being measured against — the process or method it supports or replaces |
| Reassessment trigger | Retraining, vendor model update, or drift monitoring threshold that requires a fresh review |
How to Score Risk: Influence × Consequence
Risk isn't a property of the model — it's a property of the decision the model feeds. Score two factors independently: how much the output influences the decision (sole basis for action versus one input a human weighs), and how severe the consequence would be if that output were wrong. A model scoring high on both is high risk. A model scoring high on influence but genuinely low on consequence, or vice versa, sits somewhere in the middle — and the same model can score differently depending entirely on how it's deployed.
Common Mistakes in AI Risk Assessments
- Rating the model instead of the decision. A sophisticated, highly accurate model deployed autonomously for a critical decision is higher risk than a simple model whose output a human always double-checks — accuracy and sophistication don't override decision context.
- Skipping the baseline performance reference. Without knowing what the model is being measured against, "acceptable performance" has no defensible meaning.
- Treating vendor risk documentation as sufficient. A vendor's assessment reflects general model behavior, not your specific context of use and compensating controls.
- No reassessment trigger defined. A model's risk rating from initial deployment doesn't automatically update itself when the context of use changes or the model is retrained.
How GoVal Supports AI Risk Assessment
GoVal captures the AI risk assessment as a structured record tied to each AI-embedded system, scaling required validation evidence, oversight requirements, and reassessment triggers to the documented risk score rather than applying a fixed checklist to every model. Context of use, decision influence, and consequence are recorded alongside the resulting risk tier, kept in the same audit-trailed structure as the rest of the system's validation history.
Related Topics
Frequently Asked Questions
What should an AI risk assessment for GxP systems include? +
How do you determine the risk level of an AI model? +
Is an LLM used for drafting always low risk? +
Does a vendor's own AI risk assessment satisfy the requirement? +
How often should an AI risk assessment be revisited? +
How does GoVal support AI risk assessment for GxP systems? +
Score AI risk by decision consequence, not model hype
Context of use, risk scoring, and reassessment triggers — captured as one structured record, in GoVal.
