用概率化注意力提升图文分类的准确与鲁棒性
PARIC: Probabilistic Attention Regularization for Language Guided Image Classification from Pre-trained Vison Language Models
- 引入概率注意力机制,生成带不确定性的参考注意力图
- 在多个数据集上提升分类准确率并降低模型偏差
- 适合关注可解释性与模型鲁棒性的视觉语言研究者
语言引导的注意力框架显著提升了图像分类的可解释性与性能,但依赖预训练视觉语言模型生成的确定性嵌入来构建参考注意力图,常忽视跨模态映射的多值性与病态特性。为此,我们提出PARIC,一种基于语言规范的概率化视觉注意力引导框架。该方法使预训练视觉语言模型能够生成概率性参考注意力图,在对齐文本与视觉模态的同时,融入不确定性估计,相较确定性方法表现更优。在基准测试中的实验表明,PARIC提升了预测准确性,缓解了偏差,保证了预测一致性,并增强了在多种数据集上的鲁棒性。
原文摘要 · Abstract (English)
Language-guided attention frameworks have significantly enhanced both interpretability and performance in image classification; however, the reliance on deterministic embeddings from pre-trained vision-language foundation models to generate reference attention maps frequently overlooks the intrinsic multivaluedness and ill-posed characteristics of cross-modal mappings. To address these limitations, we introduce PARIC, a probabilistic framework for guiding visual attention via language specifications. Our approach enables pre-trained vision-language models to generate probabilistic reference attention maps, which align textual and visual modalities more effectively while incorporating uncertainty estimates, as compared to their deterministic counterparts. Experiments on benchmark test problems demonstrate that PARIC enhances prediction accuracy, mitigates bias, ensures consistent predictions, and improves robustness across various datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。