用虚拟掩码增强扩散模型,少样本下修复彩色图像更精细。
ESDiff: Encoding Strategy-inspired Diffusion Model with Few-shot Learning for Color Image Inpainting
- 用虚拟掩码在通道间互扰生成高维特征,提升表征能力。
- 仅需少量训练数据,在纹理和结构上优于现有方法。
- 适合图像修复场景,尤其缺乏大量标注数据时使用。
图像修复用于恢复图像中缺失或受损区域。传统方法主要依赖邻近像素信息重建,难以保持复杂细节与结构;而深度学习模型通常需要大量训练数据。本文提出一种受编码策略启发的少样本扩散模型,用于彩色图像修复。核心思路是引入“虚拟掩码”,通过通道间的相互扰动构建高维对象,使扩散模型能从有限样本中捕捉多样化的图像表征与细节特征。该编码策略利用通道冗余,在迭代修复过程中结合低秩方法,并融入扩散模型机制,实现精准的信息输出。实验结果表明,本方法在定量指标上超越现有技术,修复图像在纹理与结构完整性方面均有提升,结果更加精确且连贯。
原文摘要 · Abstract (English)
Image inpainting is a technique used to restore missing or damaged regions of an image. Traditional methods primarily utilize information from adjacent pixels for reconstructing missing areas, while they struggle to preserve complex details and structures. Simultaneously, models based on deep learning necessitate substantial amounts of training data. To address this challenge, an encoding strategy-inspired diffusion model with few-shot learning for color image inpainting is proposed in this paper. The main idea of this novel encoding strategy is the deployment of a "virtual mask" to construct high-dimensional objects through mutual perturbations between channels. This approach enables the diffusion model to capture diverse image representations and detailed features from limited training samples. Moreover, the encoding strategy leverages redundancy between channels, integrates with low-rank methods during iterative inpainting, and incorporates the diffusion model to achieve accurate information output. Experimental results indicate that our method exceeds current techniques in quantitative metrics, and the reconstructed images quality has been improved in aspects of texture and structural integrity, leading to more precise and coherent results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。