CCAI9003 Artificial Intelligence
AI Safety: Catastrophe, Alignment, and Governance


Course Description

Artificial Intelligence (AI) systems are rapidly improving in capabilities. This rapid advancement poses a potentially catastrophic risk to humanity. This course is an interdisciplinary introduction to AI safety. Students will learn about (i) present and future capabilities of AI systems; (ii) the potentially catastrophic risk that AI systems pose to humanity; (iii) AI alignment, which seeks to give AI systems goals that reflect human values; (iv) AI governance, which seeks to design optimal regulations and institutions for controlling AI systems; and (v) foundational ethical and philosophical challenges in designing safe AI. Emphasizing an interdisciplinary approach, the course draws on perspectives from computer science, philosophy, political science, and complex systems theory. Through discussions, case studies, and policy briefs, students will critically analyse real-world scenarios where AI safety plays a pivotal role.

Course Learning Outcomes

On completing the course, students will be able to:

  1. Communicate effectively regarding core catastrophic risks surrounding AI systems.
  2. Understand core questions related to AI alignment.
  3. Explain key proposals in AI governance.
  4. Demonstrate competence in concepts related to safety engineering and complex systems.

Offer Semester and Day of Teaching

First semester (Wed)


Study Load

Activities Number of hours
Lectures 24
Tutorials 12
Reading / Self-study 60
Assessment: Essay / Report writing 24
Assessment: Presentation (incl preparation)  12
Total: 132

Assessment: 100% coursework

Assessment Tasks Weighting
Case analysis 20
Issue papers 20
Design proposal 20
Reflective journal 10
Group presentation  20
Tutorial discussion 10

Required Reading / Viewing

Philosophy of Risk

Alignment

  • Russell, S. (2022). Human-compatible artificial intelligence. In S. Muggleton & N. Chater (Eds.), Human-Like Machine Intelligence 1 (pp. 3-22). Oxford University Press.

Existing Capabilities

  • Anthropic. (2024). Alignment faking in large language models.
  • Anthropic. (2025). Agentic Misalignment: How LLMs could be insider threats.
  • Apollo Research/OpenAI. (2024). Frontier models are capable of in-context scheming.

Beneficial AI and Machine Ethics

  • Bykvist, K. (2017). Moral uncertainty. Philosophy Compass, 12(3), e12408.
  • Conitzer, V., Freedman, R., Heitzig, J., Holliday, W. H., Jacobs, B. M., Lambert, N., … & Zwicker, W. S. (2024). Social choice should guide ai alignment in dealing with diverse human feedback. arXiv preprint arXiv:2404.10271.
  • Propublica. (2016). Machine bias.

Catastrophic Risk

  • Bales, A., D’Alessandro, W., & Kirk‐Giannini, C. D. (2024). Artificial intelligence: Arguments for catastrophic risk. Philosophy Compass, 19(2), e12964.
  • Cappelen, H., Goldstein, S., & Hawthorne, J. (2026). AI survival stories: A taxonomic analysis of AI existential risk. Philosophy of AI, 1, 1-19.

Governance

Course Co-ordinator and Teacher(s)

Course Co-ordinator Contact
Professor S.D. Goldstein
School of Humanities (Philosophy), Faculty of Arts 
Email: sgold@hku.hk
Teacher(s) Contact
Professor S.D. Goldstein
School of Humanities (Philosophy), Faculty of Arts
Email: sgold@hku.hk