Module 2 • Lesson 1545 mins

Identify and Establish Defenses for the 4 Critical AI Risk Categories

Classify the 4 core AI risks (Hallucination, Bias, Edge Cases, Misuse), prioritize via Likelihood × Severity, and establish quantitative thresholds with graceful fallbacks.

Classify 4 risks: Hallucination, Bias, Distribution Shift, Misuse
Prioritize defensive resources using the Likelihood × Severity Matrix
Define quantitative acceptance thresholds and Graceful Degradation workflows

Identify and Establish Defenses for the 4 Critical AI Risk Categories

In Lesson 8, we treated Error Tolerance as a macroscopic constraint. This lesson deconstructs the operational risk surface: In what specific failure modes will the model fail? You cannot construct effective guardrails for a risk archetype you have not explicitly classified.

Running Example: FinTrack Portal internal employee assistant - allowing staff to query corporate policies, health insurance benefits, parental leave, and travel expense reimbursements from unstructured internal documentation (RAG).

TIẾP CẬN CŨ / LỖI THỜI
Naive: 'AI might make mistakes'
Spreads defensive budget equally, leaves critical vulnerabilities exposed
CHUẨN AI PM / HIỆN ĐẠI
Likelihood × Severity Risk Matrix
Concentrates guardrails on Red Zone risks and deploys seamless fallback UX
AI RISK TRIAGELikelihood × Severity Matrix & Fallback Flow

Identify and Set Guardrails for 4 Critical AI Risk Archetypes

1. Misuse & Prompt InjectionRED ZONE: MAXIMUM GUARDRAILS
REAL-WORLD SCENARIO

Employee uses Prompt Injection to force the bot to print company-wide salaries.

GUARDRAIL ARCHITECTURE

Deterministic pre-prompt filter + PII Masking + Strict RBAC database isolation.

AI Defense Principle: Allocate guardrails via the Likelihood × Severity matrix. Eliminate hallucinations with hard similarity thresholds (85%) and graceful Human-in-the-loop fallback pathways.