arXiv:2512.21104cs.CV2025-12AAAI被引 3

无需调参的图像修复方法,让生成内容更贴合文本提示且视觉合理。

FreeInpaint: Tuning-free Prompt Alignment and Visual Rationality Enhancement in Image Inpainting

  • 通过优化初始噪声和中间潜在表示,实时提升生成一致性。
  • 在多个扩散模型上验证,显著改善提示对齐与视觉合理性。
  • 插件式设计,无需训练即可适配现有修复模型,适合快速部署。

文本引导的图像修复旨在利用用户提供的文本提示,在图像指定区域生成新内容。核心挑战在于如何精准实现生成区域与提示语义的一致性,同时保持高视觉保真度。尽管现有方法借助预训练文生图扩散模型已取得视觉逼真的结果,但仍难以兼顾提示对齐与视觉合理性。本文提出 FreeInpaint,一种即插即用、无需调参的方法,在推理阶段直接优化扩散潜空间以提升生成图像的忠实度。技术上,我们引入先验引导的噪声优化策略,通过优化初始噪声引导模型注意力聚焦于有效修复区域;并精心设计针对修复任务的复合引导目标,通过优化每一步的中间潜变量,高效引导去噪过程,从而增强提示对齐与视觉合理性。大量实验表明,FreeInpaint 在多种扩散修复模型与评估指标下均表现出优异的性能与鲁棒性。

原文摘要 · Abstract (English)

Text-guided image inpainting endeavors to generate new content within specified regions of images using textual prompts from users. The primary challenge is to accurately align the inpainted areas with the user-provided prompts while maintaining a high degree of visual fidelity. While existing inpainting methods have produced visually convincing results by leveraging the pre-trained text-to-image diffusion models, they still struggle to uphold both prompt alignment and visual rationality simultaneously. In this work, we introduce FreeInpaint, a plug-and-play tuning-free approach that directly optimizes the diffusion latents on the fly during inference to improve the faithfulness of the generated images. Technically, we introduce a prior-guided noise optimization method that steers model attention towards valid inpainting regions by optimizing the initial noise. Furthermore, we meticulously design a composite guidance objective tailored specifically for the inpainting task. This objective efficiently directs the denoising process, enhancing prompt alignment and visual rationality by optimizing intermediate latents at each step. Through extensive experiments involving various inpainting diffusion models and evaluation metrics, we demonstrate the effectiveness and robustness of our proposed FreeInpaint.

图像修复扩散模型文本生成无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。