The Data Science Discipline of the School of EECS is hosting the following guest seminar:
Mechanistic Driven AI Safety and Alignment
Speaker: Dr Usman Naseem, Macquarie University
Abstract: Large language models are becoming more capable and more widely used, yet their behaviour can still diverge from human preferences, values, and expectations, leading to responses that are unsafe, biased, culturally misaligned, or needlessly restrictive. In this talk I will first discuss our work on measuring that divergence, evaluating how current models behave across a range of tasks and cultural and social settings where the right answer depends on more than getting the facts correct. I will then turn to our efforts to improve alignment through steering and inference-time intervention, before asking a deeper question: if we can observe that a model is misaligned, can we understand why it behaves that way and control it reliably rather than patching it after the fact? This motivates our work on mechanistic-driven alignment, which looks inside the model at how alignment-related behaviours emerge across layers and how we can intervene more precisely. I will close with our vision for AI systems that are not only safer, but also interpretable, controllable, trustworthy, and responsive to the diversity of human values and contexts.
Bio: Usman Naseem is a Lecturer in the School of Computing at Macquarie University, where he leads the SocialNLP Lab. His research is in Artificial Intelligence and Natural Language Processing (NLP), focusing on the safety and alignment of large language models and NLP for social good. His current work focuses on aligning models with human preferences, values, and cultures, and using mechanistic interpretability to understand and control how aligned behaviour emerges within models. He has received nine Best Paper Awards, including at AAAI 2026, and the DAAD AINet Fellowship (2023 & 2024). He has served in senior roles at leading NLP and AI conferences, including Program Co-Chair for the Web4Good track at WebConf 2026 and Senior Area Chair for EMNLP 2025, EACL 2026, and ACL 2026, and has organised shared tasks and workshops across NLP, multimodal AI, and AI for social good. More information: https://usmaann.github.io/.
About Data Science Seminar
This seminar series is hosted by EECS Data Science.
Venue
Room 46-914, or https://uqz.zoom.us/j/82323116669