arXiv:2511.12405cs.CV2025-11被引 4

让自动驾驶在陌生环境自主决策,靠视觉语言匹配动作

VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving

  • 用冻结的视觉语言模型实现无领域调优的开放世界感知
  • 通过视觉-动作对比学习,提升未知场景下的行为泛化能力
  • 适合研究开放世界自动驾驶与多模态决策的开发者

在非结构化户外环境中实现端到端自动驾驶是极具前景但极具挑战的任务,因需具备强泛化能力。本文提出视觉-语言-动作检索(VLA-R)框架,将开放世界感知与新型视觉-动作检索范式结合。利用冻结的视觉语言模型进行开放世界检测与分割,获得多尺度、提示引导且可解释的感知特征,无需特定领域调优。Q-Former瓶颈层将细粒度视觉表示与语言对齐特征融合,打通感知与动作域。为学习可迁移驾驶行为,引入视觉-动作对比学习,对齐视觉-语言与动作嵌入,实现有效的开放世界推理与动作检索。在真实机器人平台上实验表明,该方法在未见、非结构化环境中表现出强泛化与探索能力,即使数据有限。演示视频见补充材料。

原文摘要 · Abstract (English)

Exploring open-world situations in an end-to-end manner is a promising yet challenging task due to the need for strong generalization capabilities. In particular, end-to-end autonomous driving in unstructured outdoor environments often encounters conditions that were unfamiliar during training. In this work, we present Vision-Language Action Retrieval (VLA-R), an open-world end-to-end autonomous driving (OW-E2EAD) framework that integrates open-world perception with a novel vision-action retrieval paradigm. We leverage a frozen vision-language model for open-world detection and segmentation to obtain multi-scale, prompt-guided, and interpretable perception features without domain-specific tuning. A Q-Former bottleneck aggregates fine-grained visual representations with language-aligned visual features, bridging perception and action domains. To learn transferable driving behaviors, we introduce a vision-action contrastive learning scheme that aligns vision-language and action embeddings for effective open-world reasoning and action retrieval. Our experiments on a real-world robotic platform demonstrate strong generalization and exploratory performance in unstructured, unseen environments, even with limited data. Demo videos are provided in the supplementary material.

自动驾驶多模态开放世界视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。