arXiv:2505.08889cs.GRcs.CV2025-05被引 25

在图像内在空间中实现像素级精准编辑,无需微调模型

IntrinsicEdit: Precise generative image manipulation in intrinsic space

  • 基于内在图像潜在空间,分离颜色与纹理等通道进行精确控制
  • 支持多种编辑任务,如物体增删、改色、全局光照调整,效果领先
  • 无需额外数据或微调,自动处理光照变化,适合高精度图像编辑场景

生成式扩散模型虽能实现高质量图像编辑并提供直观的提示词和语义绘图接口,但这些接口缺乏精确控制,且多数方法仅针对单一编辑任务。本文提出一种通用的生成式工作流,运行于内在图像潜在空间,可在复杂图像上实现语义化、局部化的像素级编辑。基于RGB-X扩散框架,解决身份保持与内在通道耦合的关键挑战。通过精确扩散逆过程与解耦通道操作,实现高效、精准的编辑,并自动处理全局光照变化,整个过程无需额外数据收集或模型微调。在多种任务上均达到当前最优性能,包括色彩与纹理调整、物体插入与删除、全局重光照及其组合。

原文摘要 · Abstract (English)

Generative diffusion models have advanced image editing with high-quality results and intuitive interfaces such as prompts and semantic drawing. However, these interfaces lack precise control, and the associated methods typically specialize on a single editing task. We introduce a versatile, generative workflow that operates in an intrinsic-image latent space, enabling semantic, local manipulation with pixel precision for a range of editing operations. Building atop the RGB-X diffusion framework, we address key challenges of identity preservation and intrinsic-channel entanglement. By incorporating exact diffusion inversion and disentangled channel manipulation, we enable precise, efficient editing with automatic resolution of global illumination effects -- all without additional data collection or model fine-tuning. We demonstrate state-of-the-art performance across a variety of tasks on complex images, including color and texture adjustments, object insertion and removal, global relighting, and their combinations.

图像编辑扩散模型像素级控制内在空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。