Limits and Paths Forward in AI Alignment


Project dates (estimated):

January 2026 - December 2028


Name of the PhD student:

Saimun Habib


Supervisors:

Fengxiang He – School of Informatics

Dan van der Horst – School of Geosciences

Steven Travis Waller – Civil and Environmental Engineering Department, TU Dresden

Tingting Mu – Department of Computer Science, University of Manchester


Project aims:

This project investigates the formal limits of AI value alignment in large language models and other computational systems deployed in high-stakes social contexts. Contemporary alignment methods, including Reinforcement Learning from Human Feedback, Direct Preference Optimisation, and Constitutional AI, often assume that sufficiently rich human preference data can be aggregated into a stable and ethically acceptable model policy. This project examines when that assumption fails. The goal is not to reject alignment as a research programme, but to clarify its boundary conditions. By identifying what no alignment method can guarantee under particular ethical, statistical, or behavioural constraints, the project aims to support more responsible AI governance and more precise public discussion of what alignment systems can and cannot achieve.


Disciplines and subfields engaged:

  • Social choice theory

  • Cognitive science

  • Moral psychology

  • Philosophy of technology

  • Human-computer interaction


Research Themes:

  • Ethics of Algorithms

    • Ethics of Algorithmic Decision-Making

    • Algorithmic Accountability and Responsibility

    • Algorithmic Transparency and Explainability

  • Ethics and Politics of Data

    • Ethical Data Science and Data Practice

  • Emerging Technology, Health and Flourishing

    • Emerging Tech and Democratic Flourishing


Related outputs:

Publications:

  • Habib, S., Belle, V., He, F. (2026). 'Efficient Counterfactual Reasoning in ProbLog via Single World Intervention Programs.’ Under review @ Journal of Logic and Computation.

  • Habib, S., Xiao, X., Fang, M., He, F. (2026). ‘Mission Impossible: Universal LLM Moral Alignment.’ Under review @ ACM.

Previous
Previous

Anatomy of Humanitarian Innovation: Mapping Syrian Refugee Governance in Jordan in the Age of Digital Technology

Next
Next

Development and implementation of clinically relevant, ethically grounded causal models in critical care