Limits and Paths Forward in AI Alignment
Project dates (estimated):
January 2026 - December 2028
Name of the PhD student:
Saimun Habib
Supervisors:
Fengxiang He – School of Informatics
Dan van der Horst – School of Geosciences
Steven Travis Waller – Civil and Environmental Engineering Department, TU Dresden
Tingting Mu – Department of Computer Science, University of Manchester
Project aims:
This project investigates the formal limits of AI value alignment in large language models and other computational systems deployed in high-stakes social contexts. Contemporary alignment methods, including Reinforcement Learning from Human Feedback, Direct Preference Optimisation, and Constitutional AI, often assume that sufficiently rich human preference data can be aggregated into a stable and ethically acceptable model policy. This project examines when that assumption fails. The goal is not to reject alignment as a research programme, but to clarify its boundary conditions. By identifying what no alignment method can guarantee under particular ethical, statistical, or behavioural constraints, the project aims to support more responsible AI governance and more precise public discussion of what alignment systems can and cannot achieve.
Disciplines and subfields engaged:
Social choice theory
Cognitive science
Moral psychology
Philosophy of technology
Human-computer interaction
Research Themes:
Ethics of Algorithms
Ethics of Algorithmic Decision-Making
Algorithmic Accountability and Responsibility
Algorithmic Transparency and Explainability
Ethics and Politics of Data
Ethical Data Science and Data Practice
Emerging Technology, Health and Flourishing
Emerging Tech and Democratic Flourishing
Related outputs:
Publications:
Habib, S., Belle, V., He, F. (2026). 'Efficient Counterfactual Reasoning in ProbLog via Single World Intervention Programs.’ Under review @ Journal of Logic and Computation.
Habib, S., Xiao, X., Fang, M., He, F. (2026). ‘Mission Impossible: Universal LLM Moral Alignment.’ Under review @ ACM.