Women in AI Research (WiAIR)
Does Liking Yellow Make You a School Bus Driver? Hidden Failures in LLMs, with Dr. Hila Gonen
04 March 2026 59:53 WiAIR
Listen to episode
About this episode
In this conversation, Dr. Hila Gonen (Assistant Professor at the University of British Columbia) joins us to explore the deep insights into how large language models (LLMs) leak semantic information, behave across languages, and how researchers can uncover their root causes. Dr. Gonen shares her journey in interpreting AI systems, addressing biases, and controlling model outputs for safer, fairer applications.
In this episode:
- The influence of prompt elements, like colour, on model predictions
- How semantic leakage impacts model outputs unintentionally
- The role of multilinguality and modality in model safety and behaviour
- Interventional vs. observational approaches to understanding models
- Challenges in controlling and aligning AI behavior across languages and domains
- Future directions in model interpretability, safety, and causal analysis
Key Topics:
- Color and semantic influence on language model completions
- The concept of semantic leakage and examples from real prompts
- Differences between bias, hallucination, and leakage failures
- Unintended behaviours discovered through experimentation
- The importance of model interpretability and transparency
- Roots of behaviour: training data and internal representations
- Interventional analysis as a causal tool in NLP research
- Cross-lingual and cross-modal alignment in safety detection
- Challenges in evaluating safety across languages and modalities
- Strategies for building robust controls against unseen attack types
- The future of AI research: combining performance with reliability and safety
- Ethical considerations: avoiding directions that hinder societal benefits
Resources & Links:
- Lipstick on a Pig: Debiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove Them
- Does Liking Yellow Imply Driving a School Bus? Semantic Leakage in Language Models
- Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior
- OMNIGUARD: An Efficient Approach for AI Safety Moderation Across Modalities and Languages
Connect with Dr. Hila Gonen:
- https://x.com/hila_gonen</li
More AI podcast episodes
Browse all →Want to find AI jobs?
Join thousands of AI professionals finding their next opportunity