arXiv:2504.05583cs.CV2025-04被引 6

用人类注视轨迹指导模型定位,提升视觉分类准确性。

Gaze-Guided Learning: Avoiding Shortcut Bias in Visual Classification

  • 引入人类注视时间序列数据,建模注意力精确定位过程。
  • 跨模态融合人眼线索与视觉特征,纠正定位偏差。
  • 在迁移和分布外数据上显著提升分类性能,适合高精度场景。

受人类视觉注意机制启发,深度神经网络广泛采用注意力机制以学习局部判别性特征,应对复杂视觉分类任务。然而,现有方法多关注特征表达,忽视其精确定位,常导致因捷径偏见引发误分类,尤其在迁移或分布外数据集上更为明显。人类则能借助先验物体知识快速定位并比较细粒度属性,这一能力在复杂高变场景中尤为关键。为此,本文提出 Gaze-CIFAR-10 人类注视时间序列数据集,并设计双序列注视编码器,建模人类对不同局部属性的精确注意力序列。同时,使用视觉变换器(ViT)学习图像内容的序列表示。通过跨模态融合,框架将人类注视先验与机器提取的视觉序列结合,有效修正图像特征表示中的定位错误。大量定性和定量实验表明,注视引导的认知线索显著提升分类准确率。

原文摘要 · Abstract (English)

Inspired by human visual attention, deep neural networks have widely adopted attention mechanisms to learn locally discriminative attributes for challenging visual classification tasks. However, existing approaches primarily emphasize the representation of such features while neglecting their precise localization, which often leads to misclassification caused by shortcut biases. This limitation becomes even more pronounced when models are evaluated on transfer or out-of-distribution datasets. In contrast, humans are capable of leveraging prior object knowledge to quickly localize and compare fine-grained attributes, a capability that is especially crucial in complex and high-variance classification scenarios. Motivated by this, we introduce Gaze-CIFAR-10, a human gaze time-series dataset, along with a dual-sequence gaze encoder that models the precise sequential localization of human attention on distinct local attributes. In parallel, a Vision Transformer (ViT) is employed to learn the sequential representation of image content. Through cross-modal fusion, our framework integrates human gaze priors with machine-derived visual sequences, effectively correcting inaccurate localization in image feature representations. Extensive qualitative and quantitative experiments demonstrate that gaze-guided cognitive cues significantly enhance classification accuracy.

视觉分类注意力机制人机协同零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。