arXiv:2606.16960cs.CV2026-06

解决自动驾驶弱重叠摄像头的度量深度问题,提升空间一致性。

SurroundNEXO: Ego-Centric Metric Bridging for Spatially Consistent Geometry in Autonomous Driving

论文配图:SurroundNEXO: Ego-Centric Metric Bridging for Spatially Consistent Geometry in Autonomous Driving
图 1 · 摘自论文原文
  • 用自车坐标系方向编码替代密集对应关系建模跨视角几何。
  • 在NuScenes等数据集上单视角误差降33.2%,跨视图一致性提升10.5%。
  • 适用于稀疏深度提示和未见过的摄像头布局,适合实际部署场景。

现代自动驾驶依赖精确的度量3D理解进行感知、重建与规划,这要求可靠的多摄像头深度预测。然而,车载环视相机阵列的外向性导致视图间视觉重叠有限,挑战了传统基于对应关系的多视角几何假设。为此,我们提出SurroundNEXO(源自西班牙语nexo,意为几何连接),一种低重叠多摄像头度量深度框架,将跨视角推理建立在自车坐标系几何而非密集视觉对应之上。SurroundNEXO首先通过自车射线位置编码赋予图像标记全局可比的自车帧视角方向,再利用稀疏激光雷达测量作为度量锚点传播绝对尺度信息,最后逐步扩展特征交互,从视图局部建模到分解式时空推理与全局融合。该设计实现了弱重叠摄像头下的度量尺度深度预测,显著提升空间一致性。在低重叠自动驾驶基准测试(包括NuScenes、Waymo和DDAD)中,SurroundNEXO相比最先进方法,单视角误差降低33.2%,跨视图一致性提升10.5%,度量重建质量提高25.6%。其对极稀疏深度提示仍具鲁棒性,并展现出强大的零样本泛化能力至未见摄像头布局。

原文摘要 · Abstract (English)

Modern autonomous driving depends on accurate metric 3D understanding for perception, reconstruction, and planning, which in turn requires reliable multi-camera depth prediction. However, the outward-facing nature of vehicle-mounted surround-view camera rigs inherently limits visual overlap across views, challenging the correspondence-based assumptions that underpin conventional multi-view geometry. To bridge this gap, we present SurroundNEXO, named after the Spanish word nexo for a geometric link, a low-overlap multi-camera metric depth framework that grounds cross-view reasoning in ego-centric geometry rather than dense visual correspondences. Instead of directly enforcing early global fusion, SurroundNEXO first assigns image tokens globally comparable ego-frame viewing directions through Ego-Ray Positional Encoding, then uses sparse LiDAR measurements as metric anchors to propagate absolute scale cues, and finally expands feature interaction progressively from view-local modeling to decomposed spatio-temporal reasoning and global integration. This design enables metric-scale depth prediction with improved spatial consistency across weakly overlapping cameras. Across low-overlap autonomous driving benchmarks, including NuScenes, Waymo and DDAD, SurroundNEXO reduces single-view error by 33.2%, improves cross-view consistency by 10.5%, and enhances metric reconstruction quality by 25.6% compared with SOTA methods. It further remains robust under extremely sparse depth prompts and exhibits strong zero-shot generalization to unseen camera layouts.

自动驾驶度量深度多视角几何稀疏激光雷达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。