用可变顺序的自回归模型提升图像生成分类效率与准确率
Revisiting Autoregressive Models for Generative Image Classification
- 采用任意顺序自回归模型,打破固定生成顺序限制
- 在多个数据集上超越扩散模型,效率提升最高达25倍
- 适合追求高效高精度分类的科研与工程人员
类条件生成模型已成为准确且鲁棒的分类器,其中扩散模型相比其他视觉生成范式(包括自回归模型)展现出明显优势。本文重新审视基于自回归(AR)的生成分类器,发现先前方法的重要局限在于依赖固定标记顺序,这为图像理解施加了过强的归纳偏置。我们观察到,单一顺序预测更依赖部分判别性线索,而对多个标记顺序求平均能提供更全面的信号。基于此洞察,我们利用近期提出的任意顺序自回归模型,估计顺序无关的预测,从而释放自回归模型的高分类潜力。所提方法在多个图像分类基准上持续优于扩散基分类器,同时效率最高提升25倍。与当前最优的自监督判别模型相比,本方法也实现了具有竞争力的分类性能——对生成式分类器而言尤为显著。
原文摘要 · Abstract (English)
Class-conditional generative models have emerged as accurate and robust classifiers, with diffusion models demonstrating clear advantages over other visual generative paradigms, including autoregressive (AR) models. In this work, we revisit visual AR-based generative classifiers and identify an important limitation of prior approaches: their reliance on a fixed token order, which imposes a restrictive inductive bias for image understanding. We observe that single-order predictions rely more on partial discriminative cues, while averaging over multiple token orders provides a more comprehensive signal. Based on this insight, we leverage recent any-order AR models to estimate order-marginalized predictions, unlocking the high classification potential of AR models. Our approach consistently outperforms diffusion-based classifiers across diverse image classification benchmarks, while being up to 25x more efficient. Compared to state-of-the-art self-supervised discriminative models, our method delivers competitive classification performance - a notable achievement for generative classifiers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。