arXiv:2509.21247cs.CVcs.AI2025-09被引 2

用自然语言引导注意力,让模型像人一样思考。

Learning to Look: Cognitive Attention Alignment with Vision-Language Models

  • 用视觉语言模型自动生成语义注意力图
  • 在复杂数据集上显著减少模型依赖捷径
  • 无需人工标注,适合大规模训练场景

卷积神经网络常依赖表面相关性‘作弊’,引发对其决策合理性的担忧。受认知科学启发,近期方法尝试通过概念监督和解释正则化引导模型注意力,但依赖人工专家标注,难以扩展。本文提出一种可扩展框架,利用视觉语言模型通过自然语言提示自动生成语义注意力图,并引入辅助损失将CNN注意力对齐至这些语言引导的注意力图,从而提升模型决策的可靠性与认知合理性,无需人工标注。在ColoredMNIST和DecoyMNIST等挑战性数据集上的实验表明,该方法在ColorMNIST上达到当前最优性能,在DecoyMNIST上与需标注基线方法相当,验证了其更强泛化能力、更低捷径依赖性,以及更符合人类直觉的注意力模式。

原文摘要 · Abstract (English)

Convolutional Neural Networks (CNNs) frequently "cheat" by exploiting superficial correlations, raising concerns about whether they make predictions for the right reasons. Inspired by cognitive science, which highlights the role of attention in robust human perception, recent methods have sought to guide model attention using concept-based supervision and explanation regularization. However, these techniques depend on labor-intensive, expert-provided annotations, limiting their scalability. We propose a scalable framework that leverages vision-language models to automatically generate semantic attention maps using natural language prompts. By introducing an auxiliary loss that aligns CNN attention with these language-guided maps, our approach promotes more reliable and cognitively plausible decision-making without manual annotation. Experiments on challenging datasets, ColoredMNIST and DecoyMNIST, show that our method achieves state-of-the-art performance on ColorMNIST and remains competitive with annotation-heavy baselines on DecoyMNIST, demonstrating improved generalization, reduced shortcut reliance, and model attention that better reflects human intuition.

注意力机制视觉语言模型可解释性认知对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。