用可见性掩码训练,让单图3D重建更精准。
Masks make discriminative models great again!
- 用优化后的3D高斯点云生成可见区域掩码,只在可见部分训练
- 在可见区域的重建质量显著优于强基线模型
- 仅训练可见部分仍能保持整体性能,适合做图像到3D提升任务
我们提出Image2GS,一种新方法,旨在从单张图像重建逼真3D场景,重点解决图像到3D的映射问题。通过将映射(从图像生成可见内容的3D模型)与补全(推测输入中不存在的内容)分离,使任务更确定,更适合判别式模型。该方法利用优化后的3D高斯点云生成可见性掩码,在训练时排除源视图不可见区域。这种掩码训练策略显著提升了可见区域的重建质量。值得注意的是,尽管仅在掩码区域训练,Image2GS在完整场景评估上仍可与基于完整目标图像训练的先进判别模型竞争。研究揭示了判别模型在拟合未见区域时的根本困难,并证明将图像到3D映射作为独立问题并采用专用技术的优势。
原文摘要 · Abstract (English)
We present Image2GS, a novel approach that addresses the challenging problem of reconstructing photorealistic 3D scenes from a single image by focusing specifically on the image-to-3D lifting component of the reconstruction process. By decoupling the lifting problem (converting an image to a 3D model representing what is visible) from the completion problem (hallucinating content not present in the input), we create a more deterministic task suitable for discriminative models. Our method employs visibility masks derived from optimized 3D Gaussian splats to exclude areas not visible from the source view during training. This masked training strategy significantly improves reconstruction quality in visible regions compared to strong baselines. Notably, despite being trained only on masked regions, Image2GS remains competitive with state-of-the-art discriminative models trained on full target images when evaluated on complete scenes. Our findings highlight the fundamental struggle discriminative models face when fitting unseen regions and demonstrate the advantages of addressing image-to-3D lifting as a distinct problem with specialized techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。