通过三维熵分析与稀疏优化,显著提升视觉自回归模型生成效率。
ToProVAR: Efficient Visual Autoregressive Modeling via Tri-Dimensional Entropy-Aware Semantic Analysis and Sparsity Optimization
- 基于注意力熵分析三维动态,精准定位可优化参数
- 在Infinity-2B/8B上实现最高3.4倍加速,质量损失极小
- 适合追求高效高质图像生成的研究者与开发者
视觉自回归(VAR)模型虽能提升生成质量,但在生成后期面临严重效率瓶颈。本文提出一种新型优化框架ToProVAR,区别于以往依赖启发式跳过的策略,首次利用注意力熵刻画模型架构中不同维度的语义投影特性,精准识别在不同标记粒度、语义范围和生成尺度下的参数动态。基于此分析,我们发现沿标记、层与生成尺度三个关键维度存在稀疏模式,并据此设计细粒度优化策略。大量实验表明,该方法在显著加速生成过程的同时,有效保持语义一致性和细节精度,在Infinity-2B与Infinity-8B模型上实现最高达3.4倍的加速,优于传统方法。代码将公开。
原文摘要 · Abstract (English)
Visual Autoregressive(VAR) models enhance generation quality but face a critical efficiency bottleneck in later stages. In this paper, we present a novel optimization framework for VAR models that fundamentally differs from prior approaches such as FastVAR and SkipVAR. Instead of relying on heuristic skipping strategies, our method leverages attention entropy to characterize the semantic projections across different dimensions of the model architecture. This enables precise identification of parameter dynamics under varying token granularity levels, semantic scopes, and generation scales. Building on this analysis, we further uncover sparsity patterns along three critical dimensions-token, layer, and scale-and propose a set of fine-grained optimization strategies tailored to these patterns. Extensive evaluation demonstrates that our approach achieves aggressive acceleration of the generation process while significantly preserving semantic fidelity and fine details, outperforming traditional methods in both efficiency and quality. Experiments on Infinity-2B and Infinity-8B models demonstrate that ToProVAR achieves up to 3.4x acceleration with minimal quality loss, effectively mitigating the issues found in prior work. Our code will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。