让视觉Transformer自动关注人类在意的区域,性能不降反升。
Cognitive Alignment At No Cost: Inducing Human Attention Biases For Interpretable Vision Transformers

- 用人类注视图微调ViT注意力权重,引导模型关注更符合人眼习惯的区域。
- 显著提升5种评估指标下的认知对齐度,且不损失图像分类性能。
- 适合追求可解释性、希望模型行为更贴近人类的AI研究者使用。
当前主流的视觉变换器(ViTs)在处理图像时与人类注意力特征差异较大。本文研究通过微调谷歌ViT-B/16的自注意力权重,使其对齐人类显著性注视图,以缩小这一认知差距。为排除通用监督信号干扰,将微调模型与随机打乱控制组对比。结果显示,微调后在五种显著性指标上均显著提升,并诱发三种典型的人类注意力偏好:逆转了原模型对大物体的偏好而倾向小物体,增强对生物体的关注,降低极端注意力熵。贝叶斯平行分析提供决定性至极强证据,表明这种认知对齐在保持原模型在ImageNet、ImageNet-C和ObjectNet等基准上的分类性能前提下实现,无性能代价。同样方法应用于ResNet-50 CNN则导致对齐度和准确率双双下降,说明ViT的模块化自注意力机制独特地支持空间优先级与表征逻辑解耦。研究证明,基于生物学先验的注意力对齐可作为免费涌现属性,提升Transformer可解释性。
原文摘要 · Abstract (English)
For state-of-the-art image understanding, Vision Transformers (ViTs) have become the standard architecture but their processing diverges substantially from human attentional characteristics. We investigate whether this cognitive gap can be shrunk by fine-tuning the self-attention weights of Google's ViT-B/16 on human saliency fixation maps. To isolate the effects of semantically relevant signals from generic human supervision, the tuned model is compared against a shuffled control. Fine-tuning significantly improved alignment across five saliency metrics and induced three hallmark human-like biases: tuning reversed the baseline's anti-human large-object bias toward small-objects, amplified the animacy preference and diminished extreme attention entropy. Bayesian parity analysis provides decisive to very-strong evidence that this cognitive alignment comes at no cost to the model's original classification performance on in- (ImageNet), corrupted (ImageNet-C) and out-of-distribution (ObjectNet) benchmarks. An equivalent procedure applied to a ResNet-50 Convolutional Neural Network (CNN) instead degraded both alignment and accuracy, suggesting that the ViT's modular self-attention mechanism is uniquely suited for dissociating spatial priority from representational logic. These findings demonstrate that biologically grounded priors can be instilled as a free emergent property of human-aligned attention, to improve transformer interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。