arXiv:2512.04733cs.CVcs.AI2025-12被引 2

让自动驾驶理解乘客情绪,让驾驶更贴心。

E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving

  • 用情绪模型捕捉语言中的情感倾向与紧迫感
  • 融合第一人称与第三人称视角提升空间认知能力
  • 在真实数据上实现最先进的情绪-行为一致性

端到端自动驾驶系统日益采用视觉-语言-动作(VLA)模型,但通常忽略乘客情绪状态,而情绪直接影响乘坐舒适度与系统接受度。我们提出开放域端到端(OD-E2E)自动驾驶,要求自动驾驶车辆能够理解自由格式的自然语言指令,推断乘客情绪,并规划物理可行轨迹。我们提出E3AD,一种情感感知的VLA框架,通过两个受认知启发的组件增强语义理解:基于连续效价-唤醒-支配(VAD)的情感模型,从语言中捕捉语气与紧迫性;以及双路径空间推理模块,融合自中心与他中心视角,实现类人空间认知。采用以一致性为导向的训练策略,结合模态预训练与基于偏好的对齐,进一步强化情感意图与驾驶行为之间的连贯性。在多个真实世界数据集上的评估显示,E3AD在视觉定位与航点规划方面表现优异,且在情感估计中达到最先进的VAD相关性。结果表明,将情感注入VLA式驾驶可实现更符合人类期望的语义定位、路径规划与反馈。

原文摘要 · Abstract (English)

End-to-end autonomous driving (AD) systems increasingly adopt vision-language-action (VLA) models, yet they typically ignore the passenger's emotional state, which is central to comfort and AD acceptance. We introduce Open-Domain End-to-End (OD-E2E) autonomous driving, where an autonomous vehicle (AV) must interpret free-form natural-language commands, infer the emotion, and plan a physically feasible trajectory. We propose E3AD, an emotion-aware VLA framework that augments semantic understanding with two cognitively inspired components: a continuous Valenc-Arousal-Dominance (VAD) emotion model that captures tone and urgency from language, and a dual-pathway spatial reasoning module that fuses egocentric and allocentric views for human-like spatial cognition. A consistency-oriented training scheme, combining modality pretraining with preference-based alignment, further enforces coherence between emotional intent and driving actions. Across real-world datasets, E3AD improves visual grounding and waypoint planning and achieves state-of-the-art (SOTA) VAD correlation for emotion estimation. These evaluation results show that injecting emotion into VLA-style driving yields more human-aligned grounding, planning, and feedback.

情感识别自动驾驶多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。