arXiv:2608.08476cs.CV2026-08

利用3D几何先验提升立体视觉中的射线证据,增强语义场景重建的可靠性。

RayLift: Lifting Complementary Ray-Wise Evidence with 3D Geometry Priors for Semantic Scene Completion

论文配图:RayLift: Lifting Complementary Ray-Wise Evidence with 3D Geometry Priors for Semantic Scene Completion
图 1 · 摘自论文原文
  • 用冻结的3D视觉模型提取几何先验,增强场景上下文理解。
  • 通过联合建模几何差异、深度置信度和空间不确定性,自适应采样表面位置。
  • 适合需要高精度3D语义重建的自动驾驶与机器人领域研究者。

基于相机的3D语义场景重建(SSC)为自动驾驶和机器人提供全面的场景理解。然而,现有方法常将立体深度估计视为确定性几何约束,导致深度不确定性和局部对应误差直接传播至体素表示中。为此,我们提出RayLift框架,以立体几何作为度量参考,同时融合互补的射线证据,自适应地恢复可靠3D结构。RayLift首先使用互补上下文编码器,从一个冻结的3D视觉基础模型中提取几何感知先验,从而丰富场景上下文。随后引入深度射线证据提升模块,联合建模几何不一致性、深度置信度和空间不确定性,自适应地采样并加权每条相机射线上的候选表面位置。最后,语义感知体素整合器通过显式建模空间支持,将所得射线证据注入体素特征。在SemanticKITTI和SSCBench-KITTI-360上的大量实验表明,RayLift取得具有竞争力的表现,并持续优于现有方法。

原文摘要 · Abstract (English)

Camera-based 3D semantic scene completion (SSC) provides comprehensive scene understanding for autonomous driving and robotics. However, existing methods often treat stereo depth estimates as deterministic geometric constraints, causing depth uncertainty and local correspondence errors to propagate directly into voxel representations. To address this issue, we propose RayLift, a framework that uses stereo geometry as a metric reference while incorporating complementary ray evidence to recover reliable 3D structures adaptively. RayLift first employs a Complementary Context Encoder that extracts geometry-aware priors from a frozen 3D vision foundation model, thereby enriching the scene context. It then introduces a Depth Ray Evidence Lifter module that jointly models geometric dissimilarity, depth confidence, and spatial uncertainty to adaptively sample and weight candidate surface locations along each camera ray. Finally, a Semantic-Aware Voxel Integrator injects the resulting ray evidence into voxel features by explicitly modeling their spatial support. Extensive experiments on SemanticKITTI and SSCBench-KITTI-360 demonstrate that RayLift achieves competitive performance and consistently outperforms existing methods.

3D重建语义分割自动驾驶几何先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。