通过频谱结构正则化,让图像分词器更懂图像频率信息。
Structured State-Space Regularization for Generation-Friendly Image Tokenization

- 基于状态空间模型的频谱特性设计新型正则化方法
- 在重建损失仅增0.3%前提下提升生成性能
- 适合关注图像生成质量的模型优化研究者
图像分词器在现代生成模型中起核心作用,其潜在空间结构直接影响生成效果。一个关键但未被充分探索的有效潜在表示特性是频谱组织性,即跨频率成分编码信息的能力。本文提出结构化状态空间正则化,一种有原则的方法来诱导潜在空间中的频谱结构。我们通过重新审视状态空间模型(SSMs)作为模拟基函数行为的系统,推导出正则化目标。这一视角揭示了SSM隐藏状态被引导捕捉频率成分,从而产生一种新正则化项,强制潜在空间捕获图像的频谱结构。实验表明,该正则化项在重建保真度仅下降0.3%的前提下,显著提升了图像分词器的生成性能。
原文摘要 · Abstract (English)
Image tokenizers play a central role in modern generative models, where the structure of the latent space critically determines the downstream generation performance. A key but underexplored property of effective latent representations is spectral organization, the ability to encode information across frequency components. In this work, we introduce structured state-space regularization, a principled approach to inducing spectral structure in latent spaces. We derive a regularization objective by revisiting state-space models (SSMs) as systems mimicking a basis function's behavior. This perspective reveals that hidden states of SSMs are induced to capture the frequency components, resulting in a novel regularizer that enforces the latent space to capture spectral structure of images. Experiments demonstrate that our regularizer improves the generative performance of image tokenizers while incurring only minimal loss in their reconstruction fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。