通过熵分析动态优化视觉自回归模型推理,实现无训练加速
Adaptive Visual Autoregressive Acceleration via Dual-Linkage Entropy Analysis
- 基于熵变化检测推理阶段拐点,自适应决定加速时机
- 按尺度和层级动态调整删减比例,保留高信息量令牌
- 无需训练,适配各类视觉生成模型的高效推理
视觉自回归建模(VAR)因大量标记序列导致计算开销巨大。现有令牌压缩方法因未能捕捉建模动态的连续演变,存在三方面局限:启发式阶段划分、非自适应调度策略、加速范围有限,导致大量加速潜力未被挖掘。由于熵变化内在反映预测不确定性演化,可作为刻画建模动态演进的合理度量。为此,我们提出NOVA——一种基于熵分析的训练无关令牌压缩加速框架。NOVA在推理过程中在线识别尺度熵增长的拐点,自适应确定加速激活尺度;通过尺度关联与层间关联比率调节,动态计算各尺度和层的差异性令牌压缩率,剔除低熵令牌,并复用前一尺度残差所生成的缓存以加速推理并维持生成质量。大量实验与分析验证了NOVA作为简单而有效的训练无关加速框架的优越性。
原文摘要 · Abstract (English)
Visual AutoRegressive modeling (VAR) suffers from substantial computational cost due to the massive token count involved. Failing to account for the continuous evolution of modeling dynamics, existing VAR token reduction methods face three key limitations: heuristic stage partition, non-adaptive schedules, and limited acceleration scope, thereby leaving significant acceleration potential untapped. Since entropy variation intrinsically reflects the transition of predictive uncertainty, it offers a principled measure to capture modeling dynamics evolution. Therefore, we propose NOVA, a training-free token reduction acceleration framework for VAR models via entropy analysis. NOVA adaptively determines the acceleration activation scale during inference by online identifying the inflection point of scale entropy growth. Through scale-linkage and layer-linkage ratio adjustment, NOVA dynamically computes distinct token reduction ratios for each scale and layer, pruning low-entropy tokens while reusing the cache derived from the residuals at the prior scale to accelerate inference and maintain generation quality. Extensive experiments and analyses validate NOVA as a simple yet effective training-free acceleration framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。