Mozhgan Nasr, Ph.D.

Visiting Postdoc at Stanford,
Postdoc Fellow at UWaterloo

I'm

Mozhgan Nasr
About Me

About Me

Hello and welcome! I am currently a Visiting Postdoctoral Scholar at Stanford working with Prof. Marco Pavone at Autonomous Systems Lab (ASL). I am also a Postdoctoral fellow at the University of Waterloo working with Prof. Krzysztof Czarnecki at the Waterloo Intelligent Systems Engineering Lab (WISE). I obtained my Ph.D. degree in Computer Science from University of Ottawa, working under the guidance of Prof. Azzedine Boukerche at the PARADISE Laboratory. My research spans multimodal learning, large language models, and autonomous systems, with a focus on trustworthy and efficient multimodal reasoning.

University of Waterloo, Stanford
mnasraza AT gmail DOT com
News

News

Research

Research

The rapid emergence of multimodal foundation models is fundamentally changing how intelligent systems perceive, reason, and interact with the world. Models that jointly understand vision, language, and action are becoming the core intelligence behind autonomous vehicles, robots, and embodied agents. As these systems transition from research laboratories to real-world deployment, the central challenge is no longer simply improving benchmark performance—it is ensuring that they reason from visual evidence, understand uncertainty, interact naturally with humans, and make reliable decisions in complex and dynamic environments.

Recent advances in large-scale multimodal learning have demonstrated remarkable capabilities in perception, reasoning, and instruction following. However, today's models remain limited by hallucinations, weak grounding, poor interpretability, and an inability to reliably connect perception with downstream decision-making. Addressing these limitations requires new algorithmic and representational foundations that integrate perception, reasoning, memory, and action while maintaining robustness and transparency.

The ultimate goal of my research is to develop trustworthy, interactive, and scalable multimodal AI systems that integrate vision, language, and action to enable grounded reasoning, predictive world modeling, and reliable decision-making in real-world environments. To this end, I develop methods spanning multimodal foundation models, computer vision, machine learning, robotics, reinforcement learning, and explainable AI to build autonomous systems that can perceive, understand, reason, and act reliably.

My earlier work focused on driving behavior analysis and intelligent transportation systems, where I developed deep learning and inverse reinforcement learning approaches for multimodal motion forecasting, scalable driver representation learning, and graph-based models for multi-agent interaction. These efforts established the foundations for my current research on grounded multimodal reasoning, trustworthy AI, world models, and embodied intelligence.

For the full list of my publications, visit my Google scholar page.


Selected Publication

VISTAQA: Benchmarking Joint Visual Question Answering and Pixel-Level Evidence

M.N. Azadani, Yimu Wang, Yongpeng Zhu, Lihong Chen, Milan Ganai, Sean Sedwards, Marco Pavone, Krzysztof Czarnecki
Under review, 2026

We present VISTAQA, the first benchmark to jointly evaluate whether multimodal large language models not only produce the correct answer but also identify the visual evidence supporting their reasoning. Spanning six domains and diverse reasoning tasks, it reveals a substantial gap between answer accuracy and visual grounding, highlighting the need for more faithful and trustworthy multimodal AI.

leo

Rethinking the Mixture of Vision Encoders Paradigm for Enhanced Visual Understanding in Multimodal LLMs

M.N. Azadani, James Riddell, Sean Sedwards, Krzysztof Czarnecki
TMLR 2026 (J2C certificate)

LEO introduces a novel MLLM that combines multiple vision encoders through an efficient dual-branch MoVE architecture. By integrating adaptive tiling with a post-adaptation fusion strategy, LEO significantly improves visual understanding across a wide range of vision-language benchmarks while remaining effective for specialized applications such as autonomous driving, demonstrating the benefits of leveraging diverse visual representations within a unified framework.

direct

DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?

Jadelynn Dao*, Milan Ganai*, Yasmina Abukhadra, Ajay Sridhar, M.N. Azadani, Katie Luo, Clark Barrett, Jiajun Wu, Chelsea Finn, Marco Pavone
Under review, 2026

Naively scaling test-time compute is not always the answer!! More reasoning, larger models, and longer memory each may help in different situations, but often at a significant cost. DIRECT takes a step in that direction by dynamically selecting the cheapest planner that can still solve a given task, achieving strong performance while allocating compute only where it provides value.

HAWAII

HAWAII: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language Models

Yimu Wang, M.N. Azadani, Sean Sedwards, Krzysztof Czarnecki
NeurIPS 2025

HAWAII introduces an efficient framework for enhancing MLLMs by distilling knowledge from multiple visual experts into a single vision encoder. Through hierarchical visual knowledge transfer and adaptive teacher-specific routing, it significantly improves visual understanding while maintaining the efficiency required for real-world deployment.

leo-mini

LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Experts

Yimu Wang*, M.N. Azadani*, Sean Sedwards, Krzysztof Czarnecki
EMNLP 2025

LEO-MINI presents an efficient MLLM that reduces the computational cost of visual reasoning without sacrificing performance. By combining conditional visual token reduction with a mixture of multimodal experts, it significantly improves efficiency while achieving state-of-the-art results across diverse vision-language benchmarks, enabling scalable multimodal understanding for real-world applications.

htmf

Hierarchical Transformers for Motion Forecasting based on Inverse Reinforcement Learning

M.N. Azadani and A. Boukerche
IEEE transaction on Vehicular Technology, 2024

HTMF combines hierarchical transformers and inverse reinforcement learning to forecast realistic, multimodal trajectories for autonomous driving. By learning from expert human behavior and incorporating scene context through occupancy maps, it generates human-like, scene-compliant motion predictions that outperform existing approaches across diverse traffic scenarios.

capha

CAPHA: Context-Aware Path Prediction of Heterogeneous Agents

M.N. Azadani and A. Boukerche
IEEE transaction on Vehicular Technology, 2023

CAPHA presents a context-aware trajectory forecasting framework that leverages inverse reinforcement learning to predict realistic, multimodal future paths for vehicles, pedestrians, and cyclists. By integrating scene understanding with transformer-based motion prediction, it generates diverse and scene-compliant trajectories, achieving state-of-the-art performance on challenging autonomous driving benchmarks.

STAG

STAG: A novel interaction-aware path prediction method based on Spatio-Temporal Attention Graphs

M.N. Azadani and A. Boukerche
Adhoc Networks, 2023

STAG introduces a spatio-temporal graph framework for vehicle trajectory prediction that explicitly models asymmetric interactions among traffic participants. By combining graph neural networks, attention mechanisms, and directed interaction graphs, it captures complex multi-agent behaviors and achieves state-of-the-art prediction accuracy across highways, intersections, and roundabouts.

CV

CV

To view my full CV please click Here (07/2026)

Academic Services

Academic Services

Area Chair

  • NeurIPS 2026

Workshop Organizing Committee Member

  • I actively serve as a lead and co-organizer of workshops at premier ML, NLP, and computer vision conferences, including VLM4RWD at NeurIPS 2025-2026, WMEAI at ECCV 2026, GroundLM at EMNLP 2026,

Journal Reviewer

  • I am an active reviewer for leading journals, including TMLR, T-VT, T-ITS, T-IV, Pattern Recognition, Nature Communications, and others.

Confernece Reviewer

  • I am an active reviewer for leading ML and CV conferences, including CVPR, ICCV, WACV, NeurIPS, ACL, and others.