arXiv:2605.31271cs.CV2026-05被引 1

用可验证的元动作提升驾驶模型的语言-行为对齐能力

DriveMA: Driving Vision-Language-Action Models with verifiable Meta-Actions

论文配图:DriveMA: Driving Vision-Language-Action Models with verifiable Meta-Actions
图 1 · 摘自论文原文
  • 设计可验证的元动作,将未来行驶意图压缩为语言指令
  • 在Waymo数据集上达8.079分,2B与4B模型均领先
  • 适合关注语言驱动自动驾驶规划的研究者

驾驶视觉-语言-动作模型(Driving VLAs)旨在通过语言提升端到端规划性能,但语言与动作之间的鸿沟限制了其潜力。本文提出DriveMA框架,基于可验证的元动作,将未来车辆运动总结为紧凑的语言意图,可通过专家轨迹和基于轨迹的标注流程构建,并通过规则投影验证生成轨迹的一致性。该框架利用可验证性,采用以动作为中心的监督训练和数据高效的逐轮信用分配强化学习,通过密集奖励与精准信用分配显式对齐高层决策与底层轨迹规划。在基于视觉的Waymo开放数据集端到端驾驶任务中,使用2B模型取得8.060分的评分,4B模型进一步提升至8.079分;在NAV SIM上也展现出竞争力的闭环规划表现。结果表明,仅通过简单元动作接口,只要具备可验证性和语言-动作对齐优化,即可实现顶尖规划性能。代码、数据与模型已公开于https://tsinghua-mars-lab.github.io/DriveMA。

原文摘要 · Abstract (English)

Driving Vision-Language-Action Models (Driving VLAs) aim to use language to improve end-to-end planning, but the language-action gap limits this promise. We propose DriveMA, a Driving VLA framework built on verifiable meta-actions, which summarize future ego motion into compact language-domain intentions and can be constructed from expert trajectories with a trajectory-grounded annotation pipeline and can be verified against generated trajectories through rule-based projection. DriveMA exploits this verifiability with action-centric supervised training and a data-efficient turn-level credit assignment reinforcement learning framework, explicitly aligning high-level decisions with low-level trajectory planning through dense rewards and precise credit assignment. DriveMA sets a new state of the art on the Waymo Open Dataset Vision-based E2E Driving, achieving a Rater Feedback Score of 8.060 with a 2B model and further improving it to 8.079 with a 4B model; it also obtains competitive closed-loop planning performance on NAVSIM. These results show that even a simple meta-action interface can achieve state-of-the-art planning when made verifiable and optimized for language-action alignment. Code, data, and models are available at https://tsinghua-mars-lab.github.io/DriveMA.

自动驾驶语言对齐可验证性强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。