无需训练,用类别引导实现精准图像修复。
GuidPaint: Class-Guided Image Inpainting with Diffusion Models
- 通过分类器引导控制修复过程中的中间生成内容。
- 在不重新训练的前提下,显著提升修复的语义一致性与视觉真实感。
- 支持随机与确定性采样,用户可灵活选择并精修结果。
近年来,扩散模型因其强大的生成能力被广泛用于图像修复任务,取得了显著成果。基于扩散模型的多模态修复方法通常需要架构修改和重训练,计算成本高。相比之下,上下文感知的扩散修复方法利用模型固有先验调整中间去噪步骤,无需额外训练即可实现高质量修复,大幅降低计算开销。然而,这些方法对掩码区域的细粒度控制不足,常导致语义不一致或视觉不合理的内容。为此,我们提出 GuidPaint——一种无需训练、类别的图像修复框架。通过在去噪过程中引入分类器引导,GuidPaint 能精确控制掩码区域内中间生成内容,确保语义一致性和视觉真实感。此外,该框架融合随机与确定性采样策略,使用户可选择偏好中间结果并进行确定性优化。实验表明,GuidPaint 在定性与定量评估中均优于现有上下文感知修复方法。
原文摘要 · Abstract (English)
In recent years, diffusion models have been widely adopted for image inpainting tasks due to their powerful generative capabilities, achieving impressive results. Existing multimodal inpainting methods based on diffusion models often require architectural modifications and retraining, resulting in high computational cost. In contrast, context-aware diffusion inpainting methods leverage the model's inherent priors to adjust intermediate denoising steps, enabling high-quality inpainting without additional training and significantly reducing computation. However, these methods lack fine-grained control over the masked regions, often leading to semantically inconsistent or visually implausible content. To address this issue, we propose GuidPaint, a training-free, class-guided image inpainting framework. By incorporating classifier guidance into the denoising process, GuidPaint enables precise control over intermediate generations within the masked areas, ensuring both semantic consistency and visual realism. Furthermore, it integrates stochastic and deterministic sampling, allowing users to select preferred intermediate results and deterministically refine them. Experimental results demonstrate that GuidPaint achieves clear improvements over existing context-aware inpainting methods in both qualitative and quantitative evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。