arXiv:2603.25741cs.CVcs.AI2026-03被引 5

让汽车听懂自然语言指令,实现个性化驾驶规划。

Vega: Learning to Drive with Natural Language Instructions

  • 融合视觉、语言与动作的统一模型,支持多模态交互。
  • 在10万级场景数据上训练,指令跟随准确率显著提升。
  • 适合自动驾驶系统研发者和智能交通研究者参考。

视觉-语言-动作模型已重塑自动驾驶,将语言融入决策流程。然而,现有方法仅用语言描述场景或进行推理,难以响应多样化用户指令以实现个性化驾驶。为此,我们构建了一个包含约10万场景的大规模驾驶数据集InstructScene,每个场景配有多种驾驶指令及对应轨迹。随后提出统一的视觉-语言-世界-动作模型Vega,采用自回归处理视觉输入与语言指令,用扩散模型生成未来预测与轨迹。通过联合注意力促进模态间交互,并为不同模态设计独立投影层以增强能力。大量实验表明,该方法不仅规划性能优越,且具备强大的指令遵循能力,为更智能、个性化的驾驶系统铺平道路。

原文摘要 · Abstract (English)

Vision-language-action models have reshaped autonomous driving to incorporate languages into the decision-making process. However, most existing pipelines only utilize the language modality for scene descriptions or reasoning and lack the flexibility to follow diverse user instructions for personalized driving. To address this, we first construct a large-scale driving dataset (InstructScene) containing around 100,000 scenes annotated with diverse driving instructions with the corresponding trajectories. We then propose a unified Vision-Language-World-Action model, Vega, for instruction-based generation and planning. We employ the autoregressive paradigm to process visual inputs (vision) and language instructions (language) and the diffusion paradigm to generate future predictions (world modeling) and trajectories (action). We perform joint attention to enable interactions between the modalities and use individual projection layers for different modalities for more capabilities. Extensive experiments demonstrate that our method not only achieves superior planning performance but also exhibits strong instruction-following abilities, paving the way for more intelligent and personalized driving systems.

自动驾驶语言指令多模态生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。