arXiv:2601.12142eess.AScs.MM2026-01中稿 · IV被引 1

让自动驾驶听懂用户语音中的情绪,实时调整驾驶行为。

Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving

  • 将摄像头与实时语音指令联动,实现动态意图感知。
  • 相比纯视觉模型,轨迹误差降低59.4%,碰撞率下降74.4%。
  • 能识别语调、节奏等情感线索,支持情绪化驾驶决策。

视觉语言动作(VLA)模型有望提供开放词汇接口,将感知模糊性转化为语义明确的驾驶决策,但现有方法仍把语言视为推理时固定的先验。这导致模型需仅从图像中推断不断变化的目标,造成反应延迟或过度保守。我们提出EchoVLA,一种具备用户感知能力的VLA,通过耦合摄像头流与现场语音指令,实现用户意图的在线干预。我们在nuScenes数据集上添加了时间对齐、意图特定的语音指令,这些指令由自车运动描述生成合成音频。进一步地,我们构建了包含情感语音与对应轨迹的多模态思维链(CoT),用于微调基于Qwen2.5-Omni的多模态大模型。通过合成不同情绪类型语音与相应驾驶行为配对,利用语调、音高和语速中的情感线索,反映用户状态如急切或迟疑,使EchoVLA不仅能理解语音语义,还能解析情感背景,实现更细腻的情绪自适应驾驶。在开环测试中,本方法相比纯视觉基线,平均L2误差减少59.4%,碰撞率降低74.4%。nuScenes上的实验验证,EchoVLA不仅可通过语音指令引导轨迹,还能根据用户语音中检测到的情绪调节驾驶行为。

原文摘要 · Abstract (English)

Vision Language Action (VLA) models promise an open-vocabulary interface that can translate perceptual ambiguity into semantically grounded driving decisions, yet they still treat language as a static prior fixed at inference time. As a result, the model must infer continuously shifting objectives from pixels alone, yielding delayed or overly conservative maneuvers. We argue that effective VLAs for autonomous driving need an online channel in which users can influence driving with specific intentions. To this end, we present EchoVLA, a user-aware VLA that couples camera streams with in situ audio instructions. We augment the nuScenes dataset with temporally aligned, intent-specific speech commands generated by converting ego-motion descriptions into synthetic audios. Further, we compose emotional speech-trajectory pairs into a multimodal Chain-of-Thought (CoT) for fine-tuning a Multimodal Large Model (MLM) based on Qwen2.5-Omni. Specifically, we synthesize the audio-augmented dataset with different emotion types paired with corresponding driving behaviors, leveraging the emotional cues embedded in tone, pitch, and speech tempo to reflect varying user states, such as urgent or hesitant intentions, thus enabling our EchoVLA to interpret not only the semantic content but also the emotional context of audio commands for more nuanced and emotionally adaptive driving behavior. In open-loop benchmarks, our approach reduces the average L2 error by $59.4\%$ and the collision rate by $74.4\%$ compared to the baseline of vision-only perception. More experiments on nuScenes dataset validate that EchoVLA not only steers the trajectory through audio instructions, but also modulates driving behavior in response to the emotions detected in the user's speech.

自动驾驶多模态情感识别语音控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。