提出统一3D感知框架,让多摄像头驾驶系统更准地理解真实几何结构。
Geometry-Grounded Unified 3D Perception for Autonomous Driving

- 用视觉重建潜空间+视角/时间注意力,捕捉多视角时空对应关系
- 注入标定信息编码,实现带尺度的精确3D结构建模
- 一套模型同时完成检测、语义占位和深度估计,适合复杂驾驶场景
基于摄像头的自动驾驶感知需要在同步多相机流中保持一致的度量3D结构共享表示。然而现有图像框架通常依赖语义识别预训练主干,并通过下游任务特定模块引入3D几何,导致共享表示可能无法显式保留度量几何与一致3D场景结构。本文提出几何锚定统一3D感知(GeoUP)框架,将VGGT的重建导向潜空间适配至标定的多相机驾驶场景。GeoUP将跨图像交互分解为自注意力、时序注意力和视图注意力,以捕捉结构不同的时空与跨视角对应关系。进一步注入校准感知的射线图编码,提供度量尺度与相机几何信息。由此生成的几何锚定潜空间被解码用于度量深度估计、3D目标检测和语义占用预测,分别对应同一3D场景的表面、实例和体素级输出。通过联合多任务与多数据集训练,GeoUP有效利用异构标注,并泛化至多样传感器配置与感知范围。在nuScenes、Argoverse 2、Waymo、KITTI和DDAD上的大量实验表明,GeoUP在检测、占用和深度估计上均达到当前最优性能。结果验证了几何锚定表示在统一3D驾驶感知中的有效性。
原文摘要 · Abstract (English)
Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera streams. However, existing image-based frameworks often rely on backbones pretrained for semantic recognition, and introduce 3D geometry through downstream task-specific modules. As a result, their shared representations may fail to preserve explicit metric geometry and consistent 3D scene structure. In this paper, we present a Geometry-grounded Unified 3D Perception (GeoUP) framework that adapts the reconstruction-oriented latent of VGGT to calibrated, streaming multi-camera driving scenes. GeoUP factorizes cross-image interaction into self, temporal, and view attention to capture structurally distinct temporal and cross-view correspondences. It further injects calibration-aware raymap encodings to provide metric scale and camera geometry. The resulting geometry-grounded latent is decoded for metric depth estimation, 3D object detection, and semantic occupancy prediction, corresponding to surface-, instance-, and volume-level readouts of the same 3D scene. Through joint multi-task and multi-dataset training, GeoUP effectively leverages heterogeneous annotations and generalizes across diverse sensor configurations and perception ranges. Extensive experiments on nuScenes, Argoverse 2, Waymo, KITTI, and DDAD demonstrate that GeoUP achieves SOTA performance across detection, occupancy, and depth estimation. These results validate the effectiveness of geometry-grounded representations for unified 3D driving perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。