arXiv:2510.12060cs.LGcs.AI2025-10被引 2

用视觉自回归模型做分类,又快又可解释。

Your VAR Model is Secretly an Efficient and Explainable Generative Classifier

  • 基于视觉自回归模型构建生成式分类器,避免扩散模型的高算力开销。
  • 新方法在准确率与推理速度上平衡更优,实测比现有方法快2.3倍。
  • 能通过词元级互信息实现可视化解释,适合需要透明决策的场景。

生成式分类器利用条件生成模型进行分类,近年来展现出对分布偏移的鲁棒性等优良特性。然而,该领域进展主要依赖于扩散模型,其高昂的计算成本严重限制了可扩展性。这一对扩散模型的过度依赖也制约了我们对生成式分类器的理解。本文提出一种基于视觉自回归(VAR)建模最新进展的新颖生成式分类器,为研究生成式分类提供了新视角。为进一步提升性能,我们引入自适应VAR分类器+(A-VARC⁺),在准确率与推理速度间实现更优权衡,显著提升实用性。此外,我们发现基于VAR的方法与扩散模型具有根本差异:由于似然函数可解析计算,该方法可通过词元级互信息实现视觉可解释性,并在类别增量学习任务中天然抵抗灾难性遗忘。

原文摘要 · Abstract (English)

Generative classifiers, which leverage conditional generative models for classification, have recently demonstrated desirable properties such as robustness to distribution shifts. However, recent progress in this area has been largely driven by diffusion-based models, whose substantial computational cost severely limits scalability. This exclusive focus on diffusion-based methods has also constrained our understanding of generative classifiers. In this work, we propose a novel generative classifier built on recent advances in visual autoregressive (VAR) modeling, which offers a new perspective for studying generative classifiers. To further enhance its performance, we introduce the Adaptive VAR Classifier$^+$ (A-VARC$^+$), which achieves a superior trade-off between accuracy and inference speed, thereby significantly improving practical applicability. Moreover, we show that the VAR-based method exhibits fundamentally different properties from diffusion-based methods. In particular, due to its tractable likelihood, the VAR-based classifier enables visual explainability via token-wise mutual information and demonstrates inherent resistance to catastrophic forgetting in class-incremental learning tasks.

生成模型可解释性自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。