arXiv:2507.19188cs.CV2025-07ICCV被引 6

通过分两阶段重建可见与推断不可见区域,提升单目语义场景补全精度

VisHall3D: Monocular Semantic Scene Completion from Reconstructing the Visible Regions to Hallucinating the Invisible Regions

  • 分两阶段:先重建可见区域,再推断不可见区域
  • 在SemanticKITTI和SSCBench-KITTI-360上达到顶尖性能
  • 适合自动驾驶等需要高精度场景理解的应用

本文提出VisHall3D,一种新型两阶段单目语义场景补全框架,旨在解决现有方法中特征纠缠与几何不一致的问题。该框架将场景补全任务分解为两个阶段:重建可见区域(视觉)与推断不可见区域(幻觉)。第一阶段引入可见性感知投影模块VisFrontierNet,精确追踪视觉边界并保留细粒度细节。第二阶段采用幻觉网络OcclusionMAE,通过噪声注入机制生成不可见区域的合理几何结构。通过解耦这两个阶段,VisHall3D有效缓解了特征纠缠与几何不一致问题,显著提升重建质量。在SemanticKITTI和SSCBench-KITTI-360两个挑战性基准上的大量实验验证了其有效性,性能远超此前方法,为自动驾驶等应用中的更精准可靠场景理解开辟了新路径。

原文摘要 · Abstract (English)

This paper introduces VisHall3D, a novel two-stage framework for monocular semantic scene completion that aims to address the issues of feature entanglement and geometric inconsistency prevalent in existing methods. VisHall3D decomposes the scene completion task into two stages: reconstructing the visible regions (vision) and inferring the invisible regions (hallucination). In the first stage, VisFrontierNet, a visibility-aware projection module, is introduced to accurately trace the visual frontier while preserving fine-grained details. In the second stage, OcclusionMAE, a hallucination network, is employed to generate plausible geometries for the invisible regions using a noise injection mechanism. By decoupling scene completion into these two distinct stages, VisHall3D effectively mitigates feature entanglement and geometric inconsistency, leading to significantly improved reconstruction quality. The effectiveness of VisHall3D is validated through extensive experiments on two challenging benchmarks: SemanticKITTI and SSCBench-KITTI-360. VisHall3D achieves state-of-the-art performance, outperforming previous methods by a significant margin and paves the way for more accurate and reliable scene understanding in autonomous driving and other applications.

场景补全单目视觉自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。