arXiv:2509.23607cs.GRcs.CV2025-09被引 3

单图生成3D场景并可控编辑纹理,无需训练

ZeroScene: A Zero-Shot Framework for 3D Scene Generation from a Single Image and Controllable Texture Editing

  • 利用大模型先验实现零样本3D重建与纹理编辑
  • 通过点云优化和多视角损失提升场景一致性
  • 支持文本驱动的纹理编辑,效果逼真

在3D内容生成领域,单图场景重建方法难以同时保证个体资产质量和整体场景连贯性,而纹理编辑技术常无法兼顾局部连续性与多视图一致性。本文提出新框架ZeroScene,借助大视觉模型的先验知识,实现零样本的单图到3D场景重建与纹理编辑。ZeroScene从输入图像中提取物体级2D分割和深度信息,推断场景内空间关系;通过联合优化点云的3D与2D投影损失,更新物体姿态以实现精准对齐,最终构建包含前景与背景的完整连贯3D场景。此外,支持场景内物体的纹理编辑:通过约束扩散模型并引入掩码引导的渐进式图像生成策略,有效保持多视角纹理一致性,并结合基于物理的渲染(PBR)材质估计进一步提升渲染真实感。实验表明,本框架不仅确保生成资产的几何与外观准确性,还能忠实重建场景布局,并生成高度细节化、符合文本提示的纹理。

原文摘要 · Abstract (English)

In the field of 3D content generation, single image scene reconstruction methods still struggle to simultaneously ensure the quality of individual assets and the coherence of the overall scene in complex environments, while texture editing techniques often fail to maintain both local continuity and multi-view consistency. In this paper, we propose a novel system ZeroScene, which leverages the prior knowledge of large vision models to accomplish both single image-to-3D scene reconstruction and texture editing in a zero-shot manner. ZeroScene extracts object-level 2D segmentation and depth information from input images to infer spatial relationships within the scene. It then jointly optimizes 3D and 2D projection losses of the point cloud to update object poses for precise scene alignment, ultimately constructing a coherent and complete 3D scene that encompasses both foreground and background. Moreover, ZeroScene supports texture editing of objects in the scene. By imposing constraints on the diffusion model and introducing a mask-guided progressive image generation strategy, we effectively maintain texture consistency across multiple viewpoints and further enhance the realism of rendered results through Physically Based Rendering (PBR) material estimation. Experimental results demonstrate that our framework not only ensures the geometric and appearance accuracy of generated assets, but also faithfully reconstructs scene layouts and produces highly detailed textures that closely align with text prompts.

3D生成零样本纹理编辑单图重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。