arXiv:2604.19710cs.CV2026-04被引 10

提出SpanVLA框架,提升自动驾驶模型推理速度与鲁棒性。

SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model

论文配图:SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model
图 1 · 摘自论文原文
  • 用流匹配策略结合历史轨迹,加速动作生成
  • 通过负样本恢复学习,提升模型对错误行为的规避能力
  • 适合需要高效可靠决策的自动驾驶研究者

视觉-语言-动作(VLA)模型为自动驾驶提供了利用世界知识和推理能力的前景,尤其在长尾场景中表现突出。然而,现有VLA模型常因自回归生成框架导致高延迟且鲁棒性不足。本文提出SpanVLA,一种端到端自动驾驶框架,融合自回归推理与流匹配动作专家。首先,引入高效桥梁机制,利用视觉-语言模型的视觉与推理引导,基于历史轨迹初始化的流匹配策略显著降低推理时间。其次,提出基于GRPO的后训练方法,使模型不仅能从正向驾驶样本中学习,还能学习如何避免典型负向行为并掌握恢复策略。我们进一步构建了mReasoning数据集,聚焦复杂推理场景与负向-恢复样本。在NAVSIM(v1和v2)上的大量实验表明,SpanVLA性能具有竞争力。定性结果也验证了模型在多样化场景下的规划性能与鲁棒性。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models offer a promising autonomous driving paradigm for leveraging world knowledge and reasoning capabilities, especially in long-tail scenarios. However, existing VLA models often struggle with the high latency in action generation using an autoregressive generation framework and exhibit limited robustness. In this paper, we propose SpanVLA, a novel end-to-end autonomous driving framework, integrating an autoregressive reasoning and a flow-matching action expert. First, SpanVLA introduces an efficient bridge to leverage the vision and reasoning guidance of VLM to efficiently plan future trajectories using a flow-matching policy conditioned on historical trajectory initialization, which significantly reduces inference time. Second, to further improve the performance and robustness of the SpanVLA model, we propose a GRPO-based post-training method to enable the VLA model not only to learn from positive driving samples but also to learn how to avoid the typical negative behaviors and learn recovery behaviors. We further introduce mReasoning, a new real-world driving reasoning dataset, focusing on complex, reasoning-demanding scenarios and negative-recovery samples. Extensive experiments on the NAVSIM (v1 and v2) demonstrate the competitive performance of the SpanVLA model. Additionally, the qualitative results across diverse scenarios highlight the planning performance and robustness of our model.

自动驾驶多模态动作生成强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。