arXiv:2505.10566cs.CV2025-05International Conf…被引 13

用3D先验提升扩散模型的图像编辑能力,支持复杂三维变换。

3D-Fixup: Advancing Photo Editing with 3D Priors

  • 基于视频数据生成训练对,利用3D模型提供空间引导
  • 实现物体平移与旋转等复杂编辑,保持对象身份一致性
  • 适合需要真实感三维编辑的研究者与开发者

尽管扩散模型在建模图像先验方面取得显著进展,但基于单张图像的3D感知图像编辑仍具挑战性。为此,我们提出3D-Fixup框架,通过学习到的3D先验指导2D图像编辑。该框架支持物体平移和3D旋转等复杂编辑场景。我们采用基于训练的方法,利用扩散模型的生成能力,借助视频数据(自然包含真实物理动态)生成训练样本对(源帧与目标帧)。不同于仅依赖单一模型推断变换,本方法引入Image-to-3D模型提供3D引导,将2D信息显式投影至3D空间以解决难题。设计了高质量3D引导的数据生成管道。实验表明,融合3D先验后,3D-Fixup能有效实现高保真、身份一致的3D感知编辑,显著提升扩散模型在真实图像操作中的应用效果。代码已公开于https://3dfixup.github.io/

原文摘要 · Abstract (English)

Despite significant advances in modeling image priors via diffusion models, 3D-aware image editing remains challenging, in part because the object is only specified via a single image. To tackle this challenge, we propose 3D-Fixup, a new framework for editing 2D images guided by learned 3D priors. The framework supports difficult editing situations such as object translation and 3D rotation. To achieve this, we leverage a training-based approach that harnesses the generative power of diffusion models. As video data naturally encodes real-world physical dynamics, we turn to video data for generating training data pairs, i.e., a source and a target frame. Rather than relying solely on a single trained model to infer transformations between source and target frames, we incorporate 3D guidance from an Image-to-3D model, which bridges this challenging task by explicitly projecting 2D information into 3D space. We design a data generation pipeline to ensure high-quality 3D guidance throughout training. Results show that by integrating these 3D priors, 3D-Fixup effectively supports complex, identity coherent 3D-aware edits, achieving high-quality results and advancing the application of diffusion models in realistic image manipulation. The code is provided at https://3dfixup.github.io/

图像编辑3D先验扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。