让自动驾驶模型同时理解视觉、语言、几何与动作,实现精准三维环境重建。
VLGA: Vision-Language-Geometry-Action Models for Autonomous Driving

- 引入几何作为第四模态,通过点云回归损失训练模型感知密集3D空间
- 在nuScenes上实现0.50米平均误差和0.18%碰撞率,刷新VLA模型纪录
- 适合追求高精度闭环驾驶与多模态协同的自动驾驶研究者
视觉-语言-动作(VLA)模型能用语言描述场景并进行推理,但难以将动作准确锚定在复杂的三维环境中。现有方法或仅注入冻结的3D模型特征而缺乏使用激励,或依赖稀疏的框与地图损失,无法提供稠密的空间信号。我们提出VLGA,首个通过重建其行驶所经稠密3D世界来监督的视觉-语言-动作模型。VLGA将几何作为与视觉、语言、动作并列的第四模态,由像素级点云回归损失驱动的专用专家监督。在挑战性的nuScenes(开环)与Bench2Drive(闭环)数据集上的大量实验表明,相比同类VLA方法,VLGA表现更优。具体而言,在nuScenes开环评估中,未使用自车状态的条件下,其平均L2误差达0.50米,3秒内碰撞率为0.18%,达到当前最优;在闭环的Bench2Drive上,取得79.08的驾驶评分,优于最强前代VLA模型0.71分,且效率与舒适度相当。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models can describe scenes and reason about them in language, yet still struggle to ground their actions in the dense 3D world around them. Existing approaches either inject features from a frozen 3D foundation model without an objective that ensures the policy uses them, or constrain geometry with sparse box and map losses that provide no dense spatial signal. We introduce VLGA, the first vision-language-action model supervised to reconstruct the dense 3D world it drives through. VLGA introduces geometry as a fourth modality alongside vision, language, and action through a dedicated expert supervised by a per-pixel pointmap regression loss against LiDAR. Extensive experiments conducted on challenging nuScenes and Bench2Drive datasets for open-loop and closed-loop evaluations, respectively, show the superiority of VLGA over counterpart VLA methods. In particular, on open-loop nuScenes, VLGA sets a new state of the art among VLA methods without ego status, with the lowest L2 (0.50\,m average) and 3-second collision rate (0.18\%). On closed-loop Bench2Drive, VLGA attains the state-of-the-art driving score of 79.08, +0.71 over the strongest prior VLA, at comparable efficiency and comfort.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。