用概率引导的结构化稀疏化提升ViT泛化能力
Likelihood-guided Regularization in Attention Based Models
- 基于变分伊辛模型,动态学习任务自适应的结构化稀疏正则
- 在CIFAR-10/100上实现稀疏条件下更高准确率和更好不确定性估计
- 适合关注模型可解释性与高效推理的视觉Transformer研究者
Transformer架构在处理结构化高维数据的分类任务中表现优异,但其性能常依赖大规模训练数据和精细正则化以防止过拟合。本文提出一种新型的、基于似然引导的变分伊辛正则化框架,应用于视觉Transformer(ViT),同时提升模型泛化能力并动态剪枝冗余参数。该方法利用贝叶斯稀疏化技术,在模型权重上施加结构化稀疏性,支持训练过程中的自适应架构搜索。与传统固定稀疏模式的dropout不同,该方法能学习任务自适应的正则化,提高效率与可解释性。我们在MNIST、Fashion-MNIST、CIFAR-10和CIFAR-100等基准数据集上评估,结果表明在稀疏复杂数据下仍具更强泛化能力,并可对权重和选择参数进行合理不确定性量化。此外,伊辛正则器使概率估计更校准,通过不确定性感知注意力机制实现结构化特征选择。结果证明,结构化贝叶斯稀疏化有效增强Transformer架构,为标准正则化提供原则性替代方案。
原文摘要 · Abstract (English)
The transformer architecture has demonstrated strong performance in classification tasks involving structured and high-dimensional data. However, its success often hinges on large- scale training data and careful regularization to prevent overfitting. In this paper, we intro- duce a novel likelihood-guided variational Ising-based regularization framework for Vision Transformers (ViTs), which simultaneously enhances model generalization and dynamically prunes redundant parameters. The proposed variational Ising-based regularization approach leverages Bayesian sparsification techniques to impose structured sparsity on model weights, allowing for adaptive architecture search during training. Unlike traditional dropout-based methods, which enforce fixed sparsity patterns, the variational Ising-based regularization method learns task-adaptive regularization, improving both efficiency and interpretability. We evaluate our approach on benchmark vision datasets, including MNIST, Fashion-MNIST, CIFAR-10, and CIFAR-100, demonstrating improved generalization under sparse, complex data and allowing for principled uncertainty quantification on both weights and selection parameters. Additionally, we show that the Ising regularizer leads to better-calibrated probability estimates and structured feature selection through uncertainty-aware attention mechanisms. Our results highlight the effectiveness of structured Bayesian sparsification in enhancing transformer-based architectures, offering a principled alternative to standard regularization techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。