用生成式框架从单图重建可编辑的3D室内场景,满足影视游戏制作需求。
3D-RE-GEN: 3D Reconstruction of Indoor Scenes with a Generative Framework
- 分步整合检测、重建与放置模型,实现物体与背景协同生成
- 通过场景级推理还原遮挡物,生成符合光照几何的真实背景
- 支持物理合理布局,适合视觉特效与游戏开发场景
近期3D场景生成技术虽能产出视觉效果佳的成果,但现有表示方式仍难以满足影视与游戏开发中对可修改纹理网格场景的需求。当前方法存在物体分割错误、空间关系不准、背景缺失等问题。本文提出3D-RE-GEN,一种组合式框架,将单张图像重建为带有纹理的3D物体与背景。通过融合各领域最先进模型,实现了单图3D场景重建的最新性能。该流程整合资产检测、重建与放置模型,使部分模型超越原设计用途。遮挡物体的恢复被视作图像编辑任务,利用生成模型在一致光照与几何约束下进行场景级推理。与现有方法不同,3D-RE-GEN生成的空间约束背景,为真实光照与仿真提供基础。为实现物理合理的布局,采用新型4-DoF可微优化对齐物体与估计地平面。3D-RE-GEN在单图3D场景重建上达到最先进水平,通过精确相机恢复与空间优化,生成连贯且可编辑的场景。
原文摘要 · Abstract (English)
Recent advances in 3D scene generation produce visually appealing output, but current representations hinder artists' workflows that require modifiable 3D textured mesh scenes for visual effects and game development. Despite significant advances, current textured mesh scene reconstruction methods are far from artist ready, suffering from incorrect object decomposition, inaccurate spatial relationships, and missing backgrounds. We present 3D-RE-GEN, a compositional framework that reconstructs a single image into textured 3D objects and a background. We show that combining state of the art models from specific domains achieves state of the art scene reconstruction performance, addressing artists' requirements. Our reconstruction pipeline integrates models for asset detection, reconstruction, and placement, pushing certain models beyond their originally intended domains. Obtaining occluded objects is treated as an image editing task with generative models to infer and reconstruct with scene level reasoning under consistent lighting and geometry. Unlike current methods, 3D-RE-GEN generates a comprehensive background that spatially constrains objects during optimization and provides a foundation for realistic lighting and simulation tasks in visual effects and games. To obtain physically realistic layouts, we employ a novel 4-DoF differentiable optimization that aligns reconstructed objects with the estimated ground plane. 3D-RE-GEN~achieves state of the art performance in single image 3D scene reconstruction, producing coherent, modifiable scenes through compositional generation guided by precise camera recovery and spatial optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。