arXiv:2602.14834cs.CVcs.AI2026-02被引 1

发现视觉模型扫描路径的中心偏差,提出新指标揭示人类眼动的真正适配区

Debiasing Central Fixation Confounds Reveals a Peripheral "Sweet Spot" for Human-like Scanpaths in Hard-Attention Vision

  • 通过控制视野范围,找到使模型扫描路径最接近人类的感知约束区间
  • 在中等视场下,模型表现优于中心固定基线且运动模式更类人
  • 提出去中心化评分法,避免误判模型行为与人类一致

人类在视觉识别中的眼动是中心视野采样与周边上下文平衡的结果。现有任务驱动的硬注意力模型常以扫描路径与人类眼动的匹配度作为评价标准,但常用指标受数据集特有中心偏倚强烈干扰,尤其在以物体为中心的数据集上。我们使用 Gaze-CIFAR-10 证明,一个简单的中心固定基线可获得令人意外的高扫描路径得分,接近许多学习型策略。这使得标准指标过于乐观,并模糊了真实行为对齐与单纯中心倾向之间的界限。我们进一步在受限视觉条件下,系统调整中央视野大小和周边上下文,发现存在一个狭窄的‘甜点’区域:仅在此范围内,模型扫描路径在去偏后仍优于中心基线,且运动统计特性时间上类人。为此,我们提出 GCS(Gaze Consistency Score)——一种去中心化的复合指标,融合运动相似性。GCS 在中等视场、兼具中央与周边视觉时,揭示出一个稳健的甜点区域,该区域无法仅从原始扫描路径指标或准确率中看出,同时揭示了当视野过大时出现的‘捷径模式’。研究讨论了在物体为中心数据集上评估主动感知的启示,以及设计能更好分离行为对齐与中心偏倚的眼动基准的重要性。

原文摘要 · Abstract (English)

Human eye movements in visual recognition reflect a balance between foveal sampling and peripheral context. Task-driven hard-attention models for vision are often evaluated by how well their scanpaths match human gaze. However, common scanpath metrics can be strongly confounded by dataset-specific center bias, especially on object-centric datasets. Using Gaze-CIFAR-10, we show that a trivial center-fixation baseline achieves surprisingly strong scanpath scores, approaching many learned policies. This makes standard metrics optimistic and blurs the distinction between genuine behavioral alignment and mere central tendency. We then analyze a hard-attention classifier under constrained vision by sweeping foveal patch size and peripheral context, revealing a peripheral sweet spot: only a narrow range of sensory constraints yields scanpaths that are simultaneously (i) above the center baseline after debiasing and (ii) temporally human-like in movement statistics. To address center bias, we propose GCS (Gaze Consistency Score), a center-debiased composite metric augmented with movement similarity. GCS uncovers a robust sweet spot at medium patch size with both foveal and peripheral vision, that is not obvious from raw scanpath metrics or accuracy alone, and also highlights a "shortcut regime" when the field-of-view becomes too large. We discuss implications for evaluating active perception on object-centric datasets and for designing gaze benchmarks that better separate behavioral alignment from center bias.

眼动建模视觉注意力去偏评估主动感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。