arXiv:2604.04439cs.LGcs.CV2026-04

通过眼动数据拆解人类玩Atari游戏时的视觉决策机制。

Estimating Central, Peripheral, and Temporal Visual Contributions to Human Decision Making in Atari Games

论文配图:Estimating Central, Peripheral, and Temporal Visual Contributions to Human Decision Making in Atari Games
图 1 · 摘自论文原文
  • 用六种信息组合训练动作预测模型,分离中心、周边和历史状态的影响。
  • 移除周边视觉信息导致准确率下降35.27%~43.90%,影响最显著。
  • 揭示了专注主导、周边主导等决策模式,适合认知科学与人机交互研究者。

我们研究动态视觉环境中不同视觉信息源对人类决策的影响。利用带有同步眼动追踪的大型Atari-HEAD游戏数据集,提出一种受控消融框架,用于逆向解析周边视觉信息、显式注视图信息以及人类行为中的历史状态信息的贡献。在20个游戏中,针对六种信息组合训练动作预测网络。结果显示,移除周边信息导致中位数准确率下降35.27%至43.90%,影响最显著;注视信息仅造成2.11%至2.76%的下降;历史状态信息影响范围更广,为1.52%至15.51%,上限可能因周边信息泄露减少而更具价值。为进一步分析,按真实动作概率对状态进行聚类,识别出以专注为主、以周边为主及更情境化的决策阶段。结果表明,人类在Atari游戏中的决策高度依赖于视线焦点之外的信息,所提框架可从行为数据中量化各信息源贡献。

原文摘要 · Abstract (English)

We study how different visual information sources contribute to human decision making in dynamic visual environments. Using Atari-HEAD, a large-scale Atari gameplay dataset with synchronized eye-tracking, we introduce a controlled ablation framework as a means to reverse-engineer the contribution of peripheral visual information, explicit gaze information in the form of gaze maps, and past-state information from human behavior. We train action-prediction networks under six settings that selectively include or exclude these information sources. Across 20 games, peripheral information shows by far the strongest contribution, with median prediction-accuracy drops in the range of 35.27-43.90% when removed. Gaze information yields smaller drops of 2.11-2.76%, while past-state information shows a broader range of 1.52-15.51%, with the upper end likely more informative due to reduced peripheral-information leakage. To complement aggregate accuracies, we cluster states by true-action probabilities assigned by the different model configurations. This analysis identifies coarse behavioral regimes, including focus-dominated, periphery-dominated, and more contextual decision situations. These results suggest that human decision making in Atari depends strongly on information beyond the current focus of gaze, while the proposed framework provides a way to estimate such information-source contributions from behavior.

视觉决策眼动追踪Atari游戏行为建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。