用真实不完整扫描数据直接训练,生成逼真3D场景。
Seen2Scene: Completing Realistic 3D Scenes with Visibility-Guided Flow
- 基于流匹配,用可见性掩码处理真实扫描中的未知区域。
- 在复杂杂乱环境中完成场景,生成质量超越基线方法。
- 支持文本、部分扫描等多模态输入,适合真实场景重建。
我们提出Seen2Scene,首个基于流匹配、直接在不完整真实3D扫描上训练的场景补全与生成方法。不同于依赖完整合成数据的以往方法,本工作引入可见性引导的流匹配,显式掩码真实扫描中的未知区域,实现从真实世界部分观测中有效学习。我们采用截断有符号距离场(TSDF)体积表示3D场景,并通过稀疏网格编码;使用稀疏变换器高效建模复杂结构,同时对未知区域进行掩码。以3D布局框作为输入条件信号,模型可灵活适配文本或部分扫描等其他输入。通过直接从真实不完整3D扫描中学习,Seen2Scene实现了复杂杂乱真实环境下的逼真3D场景补全。实验表明,该模型生成的场景具有一致性、完整性与真实性,其补全准确率和生成质量均优于基线方法。
原文摘要 · Abstract (English)
We present Seen2Scene, the first flow matching-based approach that trains directly on incomplete, real-world 3D scans for scene completion and generation. Unlike prior methods that rely on complete and hence synthetic 3D data, our approach introduces visibility-guided flow matching, which explicitly masks out unknown regions in real scans, enabling effective learning from real-world, partial observations. We represent 3D scenes using truncated signed distance field (TSDF) volumes encoded in sparse grids and employ a sparse transformer to efficiently model complex scene structures while masking unknown regions. We employ 3D layout boxes as an input conditioning signal, and our approach is flexibly adapted to various other inputs such as text or partial scans. By learning directly from real-world, incomplete 3D scans, Seen2Scene enables realistic 3D scene completion for complex, cluttered real environments. Experiments demonstrate that our model produces coherent, complete, and realistic 3D scenes, outperforming baselines in completion accuracy and generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。