用多视角特征对齐实现更准的房间布局估计
PixCuboid: Room Layout Estimation from Multi-view Featuremetric Alignment
- 通过多视角深度特征对齐优化房间立方体布局
- 在新基准上显著超越现有方法,精度提升明显
- 支持单间与多间房间布局扩展,适用性强
粗粒度房间布局估计为众多下游任务提供重要几何线索。当前最先进方法多基于单视角图像,常假设为全景图。我们提出 PixCuboid,一种基于优化的立方体形房间布局估计方法,利用密集深度特征的多视角对齐。通过端到端训练优化,学习到具有大收敛域和光滑损失景观的特征图,使我们能用简单启发式初始化布局。评估中,我们基于 ScanNet++ 和 2D-3D-Semantics 构建两个新基准,包含人工验证的真值 3D 立方体。全面实验验证了本方法的有效性,并显著优于竞争方案。尽管网络仅训练于单个立方体,但优化框架的灵活性使其可轻松扩展至多房间布局,如大型公寓或办公室场景。代码与模型权重已公开于 https://github.com/ghanning/PixCuboid。
原文摘要 · Abstract (English)
Coarse room layout estimation provides important geometric cues for many downstream tasks. Current state-of-the-art methods are predominantly based on single views and often assume panoramic images. We introduce PixCuboid, an optimization-based approach for cuboid-shaped room layout estimation, which is based on multi-view alignment of dense deep features. By training with the optimization end-to-end, we learn feature maps that yield large convergence basins and smooth loss landscapes in the alignment. This allows us to initialize the room layout using simple heuristics. For the evaluation we propose two new benchmarks based on ScanNet++ and 2D-3D-Semantics, with manually verified ground truth 3D cuboids. In thorough experiments we validate our approach and significantly outperform the competition. Finally, while our network is trained with single cuboids, the flexibility of the optimization-based approach allow us to easily extend to multi-room estimation, e.g. larger apartments or offices. Code and model weights are available at https://github.com/ghanning/PixCuboid.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。