Lesson 317
AI Fairness, Bias & Safety
Bias · Fairness · Robustness · Alignment
1:00Why a 95%-accurate model can still be biased and unsafe — and the metrics, tools, and governance frameworks used to audit and improve it.
By the end, you can
- Explain why a high-accuracy model can still be biased or unsafe.
- Name the four structural sources of bias and give a concrete example of each.
- Define demographic parity, equalized odds, and calibration, and state what each requires.
- Describe the impossibility result and explain why fairness is a value choice rather than a purely technical one.
- Explain what LIME and SHAP do, and articulate their limitations.
- Describe adversarial examples and explain how a targeted perturbation works.
- Distinguish membership inference from data extraction, and explain differential privacy as a defense.
- Define reward hacking and explain why RLHF does not fully solve the alignment problem.
- Map AI system types to their EU AI Act risk tier and state what each tier requires.
Up next in AI, Machine Learning & Course Review




