从游戏数据中提取人类决策注意力模式,提升强化学习代理性能。
Revealing Human Attention Patterns from Gameplay Analysis for Reinforcement Learning
- 构建上下文任务相关注意力网络,从人类和智能体游戏数据生成注意力图。
- 人类注意力图更稀疏且与眼动数据匹配度更高,表明反映真实内部注意。
- 用人类注意力指导智能体,学习更稳定,效果优于基于眼动的基线方法。
本研究提出一种新方法,仅通过游戏数据揭示人类内在注意力模式(决策相关注意力),利用强化学习中的离线注意力技术。我们设计了上下文任务相关(CTR)注意力网络,从人类及强化学习智能体在Atari环境中的游戏数据生成注意力图。为验证人类CTR图是否反映内部注意力,我们通过定量与定性比较,将其与智能体注意力图以及基于人类眼动数据的时序整合外显注意力(TIOA)模型进行对照。结果表明,人类CTR图比智能体图更稀疏,且与TIOA图一致性更高。经定性视觉对比,我们推断其可能捕捉到了真实的内部注意力模式。进一步应用中,利用这些注意力图引导强化学习智能体,发现人类注意力引导的智能体相较基线实现轻微提升且学习更稳定,显著优于基于TIOA的智能体。该工作深化了对人类与智能体注意力差异的理解,并提供了一种从行为数据中提取与验证内部注意力的新路径。
原文摘要 · Abstract (English)
This study introduces a novel method for revealing human internal attention patterns (decision-relevant attention) from gameplay data alone, leveraging offline attention techniques from reinforcement learning (RL). We propose contextualized, task-relevant (CTR) attention networks, which generate attention maps from both human and RL agent gameplay in Atari environments. To evaluate whether the human CTR maps reveal internal attention patterns, we validate our model by quantitative and qualitative comparison to the agent maps as well as to a temporally integrated overt attention (TIOA) model based on human eye-tracking data. Our results show that human CTR maps are more sparse than the agent ones and align better with the TIOA maps. Following a qualitative visual comparison we conclude that they likely capture patterns of internal attention. As a further application, we use these maps to guide RL agents, finding that human attention-guided agents achieve slightly improved and more stable learning compared to baselines, and significantly outperform TIOA-based agents. This work advances the understanding of human-agent attention differences and provides a new approach for extracting and validating internal attention patterns from behavioral data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。