不重训练也不改结构,用频谱削弱提升视觉自回归模型生成质量
Guiding Visual Autoregressive Models through Spectrum Weakening
- 在频域构造可控弱模型,通过选择性保留频谱实现引导
- 无需重训练或修改结构,在离散与连续模型上均实现高质量生成
- 适合需要快速提升生成效果且不希望改动模型的开发者
无分类器指导(CFG)已成为提升生成质量与条件对齐的实用方法。尽管近期研究探索了无条件生成的引导机制,但这些方法仍依赖于扩散模型的特定假设。本文提出一种面向视觉自回归(AR)模型的频谱削弱框架。该方法无需重训练、特定条件或架构修改,通过在频域构建可控弱模型实现引导。理论上证明可逆频谱变换能保留信息,而有选择地保留部分频谱则引入可控的信息削减。基于此,我们在内部表征的通道维度上进行频谱选择,避免了扩散模型带来的结构限制。此外,我们提出了两种频谱归一化策略以确保削弱过程中的数值稳定性。在离散与连续的AR模型上,分别采用文本或类别条件进行了大量实验。结果表明,该方法在无条件生成中实现高质量输出,同时在条件生成中保持强提示对齐能力。
原文摘要 · Abstract (English)
Classifier-free guidance (CFG) has become a widely adopted and practical approach for enhancing generation quality and improving condition alignment. Recent studies have explored guidance mechanisms for unconditional generation, yet these approaches remain fundamentally tied to assumptions specific to diffusion models. In this work, we propose a spectrum-weakening framework for visual autoregressive (AR) models. This method works without the need for re-training, specific conditions, or any architectural modifications. It achieves this by constructing a controllable weak model in the spectral domain. We theoretically show that invertible spectral transformations preserve information, while selectively retaining only a subset of spectrum introduces controlled information reduction. Based on this insight, we perform spectrum selection along the channel dimension of internal representations, which avoids the structural constraints imposed by diffusion models. We further introduce two spectrum renormalization strategies that ensures numerical stability during the weakening process. Extensive experiments were conducted on both discrete and continuous AR models, with text or class conditioning. The results demonstrate that our method enables high-quality unconditional generation while maintaining strong prompt alignment for conditional generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。