arXiv:2606.14307cs.CV2026-06中稿 · ECCV

统一3D重建与全景分割,让模型一次学会看懂空间结构和物体类别。

Pano3D: Unified 3D Reconstruction and Panoptic Segmentation

论文配图:Pano3D: Unified 3D Reconstruction and Panoptic Segmentation
图 1 · 摘自论文原文
  • 用集合式掩码解码器融合几何与语义信息,联合训练提升性能。
  • 在ScanNet、ScanNet200、ScanNet++上达到当前最优全景分割效果。
  • 适用于在线与全对全注意力架构,通用性强,适合多场景应用。

最近的3D前馈重建神经网络在无需相机参数的情况下实现了密集重建的显著进展。然而,为这些模型赋予稳健的语义理解仍是未解难题。本文提出一种统一框架,实现3D重建与3D全景分割。基于现有3D重建模型,引入基于集合的掩码解码器,并通过几何与语义损失联合训练,二者相互促进。特征初始化于几何信息,随后微调以同时捕捉几何与语义。我们验证了该方法在在线与全对全注意力重建骨干上的泛化能力。在ScanNet、ScanNet200和ScanNet++数据集上,该方法实现了3D全景分割的最新性能。消融实验表明,统一模型的联合训练使3D前馈重建网络具备全景分割能力,并带来相互提升。

原文摘要 · Abstract (English)

Recent advances in 3D feedforward reconstruction neural networks have achieved remarkable success in dense reconstruction from images without any camera parameters. Yet, equipping these models with robust semantic understanding remains an open problem. Here we introduce an approach that performs 3D reconstruction and 3D panoptic segmentation in a unified framework. We build on existing 3D reconstruction models and augment them with a set-based mask decoder. The approach is jointly trained with a geometric and semantic loss, which are shown to be mutually beneficial. More precisely, the features are initialized from the geometric information and then finetuned to capture jointly geometry and semantics. We demonstrate the generality of our approach by successfully applying our framework both to online and all-to-all attention reconstruction backbones. Our method achieves state-of-the-art performance in 3D panoptic segmentation across ScanNet, ScanNet200, and ScanNet++ datasets. Ablation studies show that such joint training of a unified model equips 3D feedforward reconstruction neural networks with panoptic segmentation and yields mutually beneficial improvements.

3D重建全景分割统一框架几何语义融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。