arXiv:2506.05883cs.CVcs.AI2025-06被引 13

基于多阶段推理的视觉语言模型,提升长尾驾驶场景决策能力

HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios

  • 采用快慢双分支架构,慢分支用视觉语言模型进行分步推理
  • 在Waymo数据集上获7.7367分,比基线高2.77%,排名第二
  • 适合需要高阶语义理解的自动驾驶系统研究者参考

我们提出豪摩视觉语言模型(HMVLM),一个端到端的驾驶框架,实现受认知启发的快-慢架构中的慢分支。快速控制器输出低级转向、油门和刹车指令,而慢速规划器——一个大型视觉语言模型——生成高级意图,如“礼让行人”或“在卡车后变道”,且不增加延迟。HMVLM引入三项改进:(1) 嵌入4秒自车运动历史的选区五视角提示;(2) 多阶段思维链(CoT)提示,强制执行场景理解→驾驶决策→轨迹推断的推理流程;(3) 基于样条的轨迹后处理,消除后期抖动与急转弯。在Waymo开放数据集上训练后,这些改进使HMVLM获得7.7367的评估者反馈得分(RFS),在2025年Waymo基于视觉的端到端驾驶挑战赛中位列第二,超越公开基线2.77%。

原文摘要 · Abstract (English)

We present HaoMo Vision-Language Model (HMVLM), an end-to-end driving framework that implements the slow branch of a cognitively inspired fast-slow architecture. A fast controller outputs low-level steering, throttle, and brake commands, while a slow planner-a large vision-language model-generates high-level intents such as "yield to pedestrian" or "merge after the truck" without compromising latency. HMVLM introduces three upgrades: (1) selective five-view prompting with an embedded 4s history of ego kinematics, (2) multi-stage chain-of-thought (CoT) prompting that enforces a Scene Understanding -> Driving Decision -> Trajectory Inference reasoning flow, and (3) spline-based trajectory post-processing that removes late-stage jitter and sharp turns. Trained on the Waymo Open Dataset, these upgrades enable HMVLM to achieve a Rater Feedback Score (RFS) of 7.7367, securing 2nd place in the 2025 Waymo Vision-based End-to-End (E2E) Driving Challenge and surpassing the public baseline by 2.77%.

自动驾驶视觉语言模型多阶段推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。