arXiv:2608.01306cs.CV2026-08

让视觉模型隐空间更易生成,通过频谱引导压缩提升图像重建质量

SPAE: Spectrally Guided Autoencoder for Pretrained Visual Latents

论文配图:SPAE: Spectrally Guided Autoencoder for Pretrained Visual Latents
图 1 · 摘自论文原文
  • 用频谱分析发现高频频段信息分散且与语义纠缠
  • 设计紧凑瓶颈层抑制高频成分,提升生成与编码隐空间对齐
  • 通道掩码策略分离语义与细节,适合需高质量生成的场景

视觉基础模型(VFM)的隐空间语义丰富,适合视觉理解。近期表示自编码器方法如RAE已证明其在图像生成中具有潜力。然而,直接建模VFM隐空间仍具挑战:DiT生成的隐空间与编码器隐空间存在频谱不匹配,尤其在高频部分。我们通过通道级频谱分析发现,这些高频成分在隐空间通道中分布稀疏且与语义信息纠缠,导致DiT难以建模。为此,我们提出SPAE,一种面向生成的隐空间适配框架。SPAE采用紧凑瓶颈层,提炼稳定语义信息的同时抑制高频成分,从而改善DiT生成隐空间与编码器隐空间的对齐。此外,引入通道级掩码策略,促进瓶颈通道中语义信息与高频细节的解耦。实验表明,SPAE在视觉理解、生成质量与重建保真度之间取得良好平衡。

原文摘要 · Abstract (English)

Latents from vision foundation models (VFMs) are semantically rich and well suited for visual understanding. Recent representation autoencoder methods such as RAE have shown that they can provide promising latent spaces for image generation. However, VFM latents remain difficult to model directly: DiT-generated latents exhibit spectral mismatch with encoder latents, especially in high-frequency components. Our channel-wise spectral analysis further reveals that these high-frequency components are diffusely distributed across latent channels and entangled with semantic information, making the latent space difficult for DiT to model. To address these challenges, we propose SPAE, latent adaptation framework for generation. Specifically, SPAE employs a compact bottleneck to distill stable semantic information while suppressing high-frequency components, thereby improving the alignment between DiT-generated latents and encoder latents. In addition, we apply a channel-wise masking strategy to promote the decoupling of semantic information and high-frequency details across bottleneck channels. Experiments show that SPAE achieves a favorable balance among visual understanding, generation quality, and reconstruction fidelity.

隐空间建模图像生成频谱分析自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。