arXiv:2603.27340cs.CV2026-03

让视觉模型在准确率与人类注视轨迹间自由权衡,提升可解释性。

EVA: Bridging Performance and Human Alignment in Hard-Attention Vision Models for Image Classification

  • 基于神经科学设计硬注意力机制,用少量视觉片段逐步聚焦图像
  • 在CIFAR-10上提升注视路径相似性,同时保持高分类准确率
  • 无需人工标注注视数据,即可生成类人注视轨迹,适合可信视觉研究

单纯优化视觉模型的分类准确率会带来对齐代价,降低类人注视路径质量并削弱可解释性。我们提出EVA,一种受神经科学启发的硬注意力机制测试平台,使性能与类人度之间的权衡显式且可调节。EVA使用小规模中心-周边表示,通过基于CNN的特征提取器采样少量序列视觉片段,并引入方差控制与自适应门控以稳定和调节注意力动态。EVA仅用标准分类目标训练,无需注视监督。在带有密集人类注视标注的CIFAR-10上,EVA在DTW、NSS等指标下显著提升注视路径对齐度,同时保持竞争力的准确率。消融实验表明,基于CNN的特征提取虽提升准确率但抑制类人度,而方差控制与门控能有效恢复类人注视轨迹,性能损失极小。我们进一步在ImageNet-100上验证了EVA的可扩展性,并在无标注的COCO-Search18上评估其注视对齐能力,结果表明模型在自然场景中无需额外训练即可生成类人注视路径。总体而言,EVA为可信、可解释的主动视觉提供了原则性框架。

原文摘要 · Abstract (English)

Optimizing vision models purely for classification accuracy can impose an alignment tax, degrading human-like scanpaths and limiting interpretability. We introduce EVA, a neuroscience-inspired hard-attention mechanistic testbed that makes the performance-human-likeness trade-off explicit and adjustable. EVA samples a small number of sequential glimpses using a minimal fovea-periphery representation with CNN-based feature extractor and integrates variance control and adaptive gating to stabilize and regulate attention dynamics. EVA is trained with the standard classification objective without gaze supervision. On CIFAR-10 with dense human gaze annotations, EVA improves scanpath alignment under established metrics such as DTW, NSS, while maintaining competitive accuracy. Ablations show that CNN-based feature extraction drives accuracy but suppresses human-likeness, whereas variance control and gating restore human-aligned trajectories with minimal performance loss. We further validate EVA's scalability on ImageNet-100 and evaluate scanpath alignment on COCO-Search18 without COCO-Search18 gaze supervision or finetuning, where EVA yields human-like scanpaths on natural scenes without additional training. Overall, EVA provides a principled framework for trustworthy, human-interpretable active vision.

视觉模型注意力机制可解释性类人行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。