arXiv:2507.15266cs.ROcs.SY2025-07被引 3

用视觉语言模型提升城市自动驾驶的决策与控制透明性

VLM-UDMC: VLM-Enhanced Unified Decision-Making and Motion Control for Urban Autonomous Driving

  • 上层慢系统用RAG增强推理,融合场景理解与风险感知
  • 实时环境变化通过上下文势函数编码,动态调整路径规划
  • 轻量级LSTM预测异构交通参与者轨迹,适合真实驾驶场景

城市自动驾驶中,场景理解与风险感知对安全决策至关重要。为模仿人类驾驶员的认知能力并保障透明性与可解释性,本文提出一种基于视觉语言模型(VLM)的统一决策与运动控制框架——VLM-UDMC。该框架将场景推理与风险感知融入上层慢系统,根据实时环境变化动态重构下层快系统的最优运动规划,环境变化通过上下文感知的势函数编码。上层系统采用两阶段检索增强生成(RAG)策略,利用基础模型处理多模态输入并检索上下文知识,生成风险感知洞察;下层则使用轻量级多核分解LSTM,通过提取更平滑的趋势表征,实现短时程异构交通参与者轨迹的实时预测。在仿真与全尺寸自动驾驶车辆的真实实验中验证了VLM-UDMC的有效性,结果表明该框架能有效结合场景理解与注意力分解,实现理性驾驶决策,显著提升城市道路整体表现。开源项目已发布于https://github.com/henryhcliu/vlmudmc.git。

原文摘要 · Abstract (English)

Scene understanding and risk-aware attentions are crucial for human drivers to make safe and effective driving decisions. To imitate this cognitive ability in urban autonomous driving while ensuring the transparency and interpretability, we propose a vision-language model (VLM)-enhanced unified decision-making and motion control framework, named VLM-UDMC. This framework incorporates scene reasoning and risk-aware insights into an upper-level slow system, which dynamically reconfigures the optimal motion planning for the downstream fast system. The reconfiguration is based on real-time environmental changes, which are encoded through context-aware potential functions. More specifically, the upper-level slow system employs a two-step reasoning policy with Retrieval-Augmented Generation (RAG), leveraging foundation models to process multimodal inputs and retrieve contextual knowledge, thereby generating risk-aware insights. Meanwhile, a lightweight multi-kernel decomposed LSTM provides real-time trajectory predictions for heterogeneous traffic participants by extracting smoother trend representations for short-horizon trajectory prediction. The effectiveness of the proposed VLM-UDMC framework is verified via both simulations and real-world experiments with a full-size autonomous vehicle. It is demonstrated that the presented VLM-UDMC effectively leverages scene understanding and attention decomposition for rational driving decisions, thus improving the overall urban driving performance. Our open-source project is available at https://github.com/henryhcliu/vlmudmc.git.

自动驾驶视觉语言模型决策控制多智能体预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。