用虚拟多视角增强单图室内语义场景重建
Fake It To Make It: Virtual Multiviews to Enhance Monocular Indoor Semantic Scene Completion
- 通过虚拟相机生成多视角图像,提升3D场景上下文信息
- 在NYUv2数据集上提升语义场景完成度4.9%的IoU
- 适合做单目3D场景理解的研究者与工程师参考
单目室内语义场景重建(SSC)旨在从一张室内场景的RGB图像中恢复3D语义占据地图,通过2D图像线索推断空间布局和物体类别。该任务的难点在于将2D图像转换为3D空间时产生的深度、尺度和形状歧义,尤其在结构复杂且频繁遮挡的室内环境中更为显著。现有方法常因这些歧义导致物体表征失真或缺失。为此,本文提出一种新方法,利用新视角合成与多视角融合技术:在场景周围设置虚拟相机,模拟多视角输入以增强上下文信息;引入多视角融合适配器(MVFA),有效整合多视角3D预测结果,生成统一的3D语义占据地图。此外,我们揭示了生成式方法在SSC中的固有局限——新颖性与一致性权衡问题。所提出的系统GenFuSE,在集成现有SSC网络后,于NYUv2数据集上实现场景完成度2.8%、语义场景完成度4.9%的IoU提升。本工作为基于合成输入推进单目SSC提供了标准框架。
原文摘要 · Abstract (English)
Monocular Indoor Semantic Scene Completion (SSC) aims to reconstruct a 3D semantic occupancy map from a single RGB image of an indoor scene, inferring spatial layout and object categories from 2D image cues. The challenge of this task arises from the depth, scale, and shape ambiguities that emerge when transforming a 2D image into 3D space, particularly within the complex and often heavily occluded environments of indoor scenes. Current SSC methods often struggle with these ambiguities, resulting in distorted or missing object representations. To overcome these limitations, we introduce an innovative approach that leverages novel view synthesis and multiview fusion. Specifically, we demonstrate how virtual cameras can be placed around the scene to emulate multiview inputs that enhance contextual scene information. We also introduce a Multiview Fusion Adaptor (MVFA) to effectively combine the multiview 3D scene predictions into a unified 3D semantic occupancy map. Finally, we identify and study the inherent limitation of generative techniques when applied to SSC, specifically the Novelty-Consistency tradeoff. Our system, GenFuSE, demonstrates IoU score improvements of up to 2.8% for Scene Completion and 4.9% for Semantic Scene Completion when integrated with existing SSC networks on the NYUv2 dataset. This work introduces GenFuSE as a standard framework for advancing monocular SSC with synthesized inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。