arXiv:2604.00813cs.CVcs.AI2026-04被引 9

用在线3D几何重建提升自动驾驶决策能力

DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale

论文配图:DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale
图 1 · 摘自论文原文
  • 提出流式处理的DVGT-2模型,实时输出3D几何与轨迹规划
  • 在nuScenes等数据集上实现更快推理速度和更优几何重建效果
  • 无需微调即可适配不同摄像头配置,适合大规模自动驾驶部署

端到端自动驾驶正从基于稀疏感知的传统范式转向视觉-语言-行动(VLA)模型,后者通过学习语言描述辅助规划。本文提出一种替代性视觉-几何-行动(VGA)范式,强调密集3D几何是自动驾驶的关键线索。鉴于车辆在三维世界中运行,我们认为密集3D几何能提供最全面的决策信息。然而,现有几何重建方法(如DVGT)依赖计算昂贵的多帧批量处理,无法用于在线规划。为此,我们提出流式驾驶视觉几何变换器DVGT-2,以在线方式处理输入,并联合输出当前帧的密集几何与轨迹规划。采用时序因果注意力和历史特征缓存支持实时推理。为提升效率,提出滑动窗口流式策略,利用特定区间内的历史缓存避免重复计算。尽管速度更快,DVGT-2在多个数据集上仍达到更优的几何重建性能。同一训练好的DVGT-2可直接用于不同相机配置下的规划,无需微调,涵盖闭环NAVSIM与开环nuScenes基准。

原文摘要 · Abstract (English)

End-to-end autonomous driving has evolved from the conventional paradigm based on sparse perception into vision-language-action (VLA) models, which focus on learning language descriptions as an auxiliary task to facilitate planning. In this paper, we propose an alternative Vision-Geometry-Action (VGA) paradigm that advocates dense 3D geometry as the critical cue for autonomous driving. As vehicles operate in a 3D world, we think dense 3D geometry provides the most comprehensive information for decision-making. However, most existing geometry reconstruction methods (e.g., DVGT) rely on computationally expensive batch processing of multi-frame inputs and cannot be applied to online planning. To address this, we introduce a streaming Driving Visual Geometry Transformer (DVGT-2), which processes inputs in an online manner and jointly outputs dense geometry and trajectory planning for the current frame. We employ temporal causal attention and cache historical features to support on-the-fly inference. To further enhance efficiency, we propose a sliding-window streaming strategy and use historical caches within a certain interval to avoid repetitive computations. Despite the faster speed, DVGT-2 achieves superior geometry reconstruction performance on various datasets. The same trained DVGT-2 can be directly applied to planning across diverse camera configurations without fine-tuning, including closed-loop NAVSIM and open-loop nuScenes benchmarks.

自动驾驶3D几何流式处理视觉-几何-行动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。