让驾驶助手懂车电状态,做出更节能的决策
EVLA: An Electro-Aware Multimodal Assistant for Physically-Grounded Driving Reasoning and Control

- 融合视觉、语言和车辆实时电控数据,生成统一表征
- 在基准测试中准确率提升5.6%,得分提高0.0871
- 适合需要物理精准控制的自动驾驶系统研发
当前驾驶辅助的视觉-语言模型通常将车辆动力学视为黑箱,导致决策缺乏对车辆实时电机械状态的感知。为此,我们提出电觉多模态助手EVLA——一个结合多模态场景理解与电动动力系统实时状态感知(如电机扭矩、电池电量)的新框架。其核心创新包括:一是统一共状态编码器(UCSE),将视觉、文本与车辆状态输入融合为共享潜在表示,并引入能量效率场以建模空间能耗;二是电觉结构化推理链(ESRC),用内部确定性推理替代外部思维链提示,基于物理约束与优化目标进行决策。通过物理引导的联合损失端到端训练,EVLA学习生成上下文感知且能效最优的驾驶决策。在驾驶问答基准上评估显示,相比强基线微调的VLM,EVLA最终得分提升+0.0871,准确率提高+5.6%。消融实验验证各组件必要性,效率分析表明其推理速度比多阶段流程快36%。本工作表明,整合车辆状态感知与结构化物理推理是构建下一代物理可信驾驶助手的关键。
原文摘要 · Abstract (English)
Modern vision-language models (VLMs) for driving assistants typically treat vehicle dynamics as a black box, resulting in decisions that lack awareness of the vehicle's real-time electro-mechanical state. To bridge this gap, we introduce the Electro-Visual-Language Assistant (EVLA) -- a novel framework that combines multi-modal scene understanding with real-time perception of the electrified powertrain state (e.g., motor torque, battery SOC). Our approach features two key innovations: first, a Unified Co-State Encoder (UCSE) that fuses visual, textual, and vehicle-state inputs into a shared latent representation, augmented with an Energy-Efficiency Field to model spatial energy costs; and second, an Electro-aware Structured Reasoning Chain (ESRC), which replaces external chain-of-thought prompting with an internal, deterministic reasoning process grounded in physical constraints and optimization objectives. Trained end-to-end with a physics-guided joint loss, EVLA learns to generate context-aware and energy-optimal driving decisions. Extensive evaluations on a driving QA benchmark demonstrate that EVLA substantially outperforms strong fine-tuned VLM baselines, improving the final score by +0.0871 and accuracy by +5.6\%. Ablation studies validate the necessity of each component, and efficiency analyses show that EVLA achieves 36\% faster inference than multi-stage pipelines. This work underscores that integrating vehicle-state awareness and structured physical reasoning is crucial for developing next-generation, physically-grounded driving assistants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。