arXiv:2511.22249cs.CV2025-11被引 6

提升高维隐空间生成质量,关键在增强高频信号训练暴露。

Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective

  • 提出频域渐进训练策略FreqWarm,增强早期高频信号学习。
  • 在多个高维自编码器上降低gFID:最高14.11,生成质量显著提升。
  • 无需修改模型即可适配多种架构,适合视觉生成研究者使用。

隐空间扩散已成为视觉生成的主流范式,但随着隐空间维度增加,重建与生成质量之间存在持续的权衡:更高容量的自编码器虽提升重建保真度,但生成质量最终下降。我们发现这一差距源于高频成分在编码和解码阶段的不同表现。通过在RGB与隐空间域进行受控扰动分析,发现解码器严重依赖高频隐变量恢复细节,而编码器对高频内容表征不足,导致扩散模型训练中高频部分暴露不够、拟合不充分。为此,我们提出即插即用的频率暖身策略FreqWarm,可在不修改或重训自编码器的前提下,增强扩散或流匹配训练初期对高频隐信号的暴露。在多个高维自编码器上应用后,生成质量持续提升:Wan2.2-VAE的gFID下降14.11,LTX-VAE下降6.13,DC-AE-f32下降4.42。研究表明,显式管理频率暴露可有效使高维隐空间更适于扩散生成。

原文摘要 · Abstract (English)

Latent diffusion has become the default paradigm for visual generation, yet we observe a persistent reconstruction-generation trade-off as latent dimensionality increases: higher-capacity autoencoders improve reconstruction fidelity but generation quality eventually declines. We trace this gap to the different behaviors in high-frequency encoding and decoding. Through controlled perturbations in both RGB and latent domains, we analyze encoder/decoder behaviors and find that decoders depend strongly on high-frequency latent components to recover details, whereas encoders under-represent high-frequency contents, yielding insufficient exposure and underfitting in high-frequency bands for diffusion model training. To address this issue, we introduce FreqWarm, a plug-and-play frequency warm-up curriculum that increases early-stage exposure to high-frequency latent signals during diffusion or flow-matching training -- without modifying or retraining the autoencoder. Applied across several high-dimensional autoencoders, FreqWarm consistently improves generation quality: decreasing gFID by 14.11 on Wan2.2-VAE, 6.13 on LTX-VAE, and 4.42 on DC-AE-f32, while remaining architecture-agnostic and compatible with diverse backbones. Our study shows that explicitly managing frequency exposure can successfully turn high-dimensional latent spaces into more diffusible targets.

隐空间生成扩散模型频率分析自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。