通过引入鸟瞰图辅助任务,提升驾驶视觉语言模型在不同摄像头配置间的零样本迁移能力。
Towards Zero-Shot Transfer Across Embodiments For Driving VLAs

- 设计鸟瞰图强制机制,让模型共享地面平面物体布局表示
- 在少量摄像头配置下,显著提升分布内与分布外性能
- 揭示数据规模扩大后辅助任务效果减弱,提醒需考虑训练多样性
视觉-语言-动作模型(VLAs)在自动驾驶中展现出强大潜力,得益于多模态预训练实现指令遵循、视觉推理和场景级泛化。在机器人操作中,跨多种机器人配置的VLA微调能提升性能和跨实体泛化;但在自动驾驶领域,VLAs仍主要针对单一数据集训练,极少评估其在未见数据集和摄像头配置下的零样本迁移能力。此外,简单增加训练数据未必提升已知配置下的表现。为此,本文研究驾驶任务的多数据集训练,并提出BEV-Forcing——一种将专用鸟瞰图模型中的地面平面物体布局信息注入VLA主干的辅助目标。通过鼓励模型通过共享的鸟瞰图空间接口表示物体位置,我们发现该辅助任务在少量摄像头配置下可同时提升分布内与分布外性能。然而,随着训练实体数量增加,辅助任务的优势逐渐减弱,表明当前文献中某些技术在单纯扩大训练多样性时效果可能下降,因此强调应结合数据扩展性分析结果。
原文摘要 · Abstract (English)
Vision-Language-Action models (VLAs) have shown strong potential in autonomous driving by leveraging multimodal pretraining for instruction following, visual reasoning, and scene-level generalization. In robotic manipulation, scaling VLA fine-tuning across multiple robot setups--especially when unifying representations across embodiments--has been shown to improve in-dataset performance and cross-embodiment generalization; in autonomous driving, however, VLAs remain largely trained on individual datasets and are rarely evaluated for zero-shot transfer to unseen datasets and camera rigs; furthermore naively adding more datasets to the training data does not necessarily lead to better performance within seen embodiments. To address these problems, we study multi-dataset training for the driving task and BEV-Forcing, an auxiliary objective that transfers ground-plane object-layout information from a specialized Bird's-Eye-View model into the VLA backbone. By encouraging the model to represent object position through a shared BEV spatial interface, we show that an auxiliary task such as BEV-Forcing can improve both in-distribution and out-of-distribution performance when training on a small number of camera rigs. As the number of training embodiments increases, however, the benefits of the auxiliary task are reduced; we present this as evidence that new techniques in the literature may see their benefits diminish when simply scaling up training diversity, which motivates presenting results taking into account data scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。