分离高低频特征,让生成图像更清晰真实
Decoupling High and Low Frequencies for Faithful Image Generation with Fine Details
- 设计解耦频段的分块编码器,分别学习高低频嵌入
- 在纹理区域生成更锐利细节,整体图像更逼真
- 适合追求高保真图像生成的研究者与应用
潜在生成模型在合成前将图像压缩为学习到的嵌入表示,生成质量关键取决于这些嵌入对视觉细节的保留程度。我们发现,尽管嵌入能有效重建低频结构,却难以恢复对感知真实感至关重要的高频细节。传统重建目标隐式优先考虑粗略结构信息,导致输出过度平滑,纹理区域质量下降。为此,我们提出 DeBaT——一种解耦频段的分块编码器,显式分离低频与高频嵌入的学习。该解耦机制在保持全局一致性的前提下,实现精细细节的准确重建。集成至基于潜在扩散的生成模型后,DeBaT 生成的样本比以往潜在编码器更锐利、更真实,验证了显式解耦高低频带有助于在嵌入空间中更好保留视觉细节。
原文摘要 · Abstract (English)
Latent generative models compress images into learned embeddings prior to synthesis, and the generation quality critically depends on how faithfully these embeddings preserve visual detail. We observe that while such embeddings are effective at reconstructing low frequency structure, they struggle to recover sharp high frequency details that are essential for perceptual realism. Conventional reconstruction objectives implicitly prioritize coarse structural information over high frequency content, which can lead to overly smoothed outputs and degraded visual quality in textured regions. Motivated by this observation, we propose DeBaT, a Decoupled frequency Band Tokenizer that explicitly separates the learning of low and high frequency band embeddings. This decoupling enables accurate reconstruction of fine details while preserving global coherence. Integrated into a latent diffusion based generative model, DeBaT allows for sharper and more realistic samples than previous latent tokenizers, confirming that the explicit decoupling of high and low frequency bands eases the preservation of visual details in learned embedding spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。