Lesson 317

AI Fairness, Bias & Safety

Bias · Fairness · Robustness · Alignment

1:00

Why a 95%-accurate model can still be biased and unsafe — and the metrics, tools, and governance frameworks used to audit and improve it.

By the end, you can

  • Explain why a high-accuracy model can still be biased or unsafe.
  • Name the four structural sources of bias and give a concrete example of each.
  • Define demographic parity, equalized odds, and calibration, and state what each requires.
  • Describe the impossibility result and explain why fairness is a value choice rather than a purely technical one.
  • Explain what LIME and SHAP do, and articulate their limitations.
  • Describe adversarial examples and explain how a targeted perturbation works.
  • Distinguish membership inference from data extraction, and explain differential privacy as a defense.
  • Define reward hacking and explain why RLHF does not fully solve the alignment problem.
  • Map AI system types to their EU AI Act risk tier and state what each tier requires.
Up next in AI, Machine Learning & Course Review
Questions or feedback?