arXiv:2503.09830cs.CV2025-03被引 4

解决扩散模型生成高清图时的重复失真问题

Exploring Position Encoding in Diffusion U-Net for Training-free High-resolution Image Generation

  • 通过动态虚拟边界增强位置信息传播
  • 无需训练即可实现高质量高清图像生成
  • 适合追求高分辨率生成效果的研究者

通过预训练U-Net去噪更高分辨率潜在表示,常导致重复和混乱的图像模式。尽管近期研究尝试通过对齐原始与高分辨率下的去噪过程来提升生成质量,但根本原因仍未充分探索。通过对U-Net中位置编码的全面分析,我们发现其根源在于位置信息在卷积层中从零填充向潜在特征传播不足,随分辨率升高而加剧。为此,我们提出一种新颖的无训练方法——渐进式边界补全(PBC),在特征图内部创建动态虚拟图像边界,以增强位置信息传播,从而实现高质量、内容丰富的高分辨率图像合成。大量实验表明该方法具有显著优势。

原文摘要 · Abstract (English)

Denoising higher-resolution latents via a pre-trained U-Net leads to repetitive and disordered image patterns. Although recent studies make efforts to improve generative quality by aligning denoising process across original and higher resolutions, the root cause of suboptimal generation is still lacking exploration. Through comprehensive analysis of position encoding in U-Net, we attribute it to inconsistent position encoding, sourced by the inadequate propagation of position information from zero-padding to latent features in convolution layers as resolution increases. To address this issue, we propose a novel training-free approach, introducing a Progressive Boundary Complement (PBC) method. This method creates dynamic virtual image boundaries inside the feature map to enhance position information propagation, enabling high-quality and rich-content high-resolution image synthesis. Extensive experiments demonstrate the superiority of our method.

扩散模型高清生成位置编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。