研究人如何理解强化学习智能体的训练过程,提升人机协作透明度。
Can you see how I learn? Human observers' inferences about Reinforcement Learning agents' learning processes
- 通过观察实验分析人类对智能体学习行为的推断机制
- 发现四类核心认知主题,且随时间动态演变
- 为可解释强化学习系统设计提供实证依据
强化学习(RL)智能体的学习行为常难以被人类直观理解,导致人机协同教学中反馈效果不佳。本文采用自下而上的两阶段实验方法,首次构建了基于观察的直接评估范式,用于探究人类对智能体学习过程的推断。在探索性访谈研究(N=9)中,识别出四类核心认知主题:智能体目标、知识状态、决策方式与学习机制。在验证性研究(N=34)中,该范式应用于导航与操作两类任务及两种算法(表格型/函数逼近),共收集816份回答。结果证实了范式的可靠性,并细化了主题演化规律及其相互关系。研究揭示了人类如何构建对智能体学习的理解,为设计可解释的强化学习系统和提升人机交互透明性提供了可操作的洞见。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) agents often exhibit learning behaviors that are not intuitively interpretable by human observers, which can result in suboptimal feedback in collaborative teaching settings. Yet, how humans perceive and interpret RL agent's learning behavior is largely unknown. In a bottom-up approach with two experiments, this work provides a data-driven understanding of the factors of human observers' understanding of the agent's learning process. A novel, observation-based paradigm to directly assess human inferences about agent learning was developed. In an exploratory interview study (\textit{N}=9), we identify four core themes in human interpretations: Agent Goals, Knowledge, Decision Making, and Learning Mechanisms. A second confirmatory study (\textit{N}=34) applied an expanded version of the paradigm across two tasks (navigation/manipulation) and two RL algorithms (tabular/function approximation). Analyses of 816 responses confirmed the reliability of the paradigm and refined the thematic framework, revealing how these themes evolve over time and interrelate. Our findings provide a human-centered understanding of how people make sense of agent learning, offering actionable insights for designing interpretable RL systems and improving transparency in Human-Robot Interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。