arXiv:2510.16956cs.AI2025-10中稿 · ECAI 2025 Workshop…被引 1

测试用户能否从XRL解释中识别智能体目标,发现仅一种算法有效。

A Comparative User Evaluation of XRL Explanations using Goal Identification

  • 设计新评测方法:通过目标识别任务评估XRL解释效果。
  • 四种算法中仅一种超越随机准确率(约50%)。
  • 用户自信程度与实际准确率无关,提示解释可读性存在误导。

可解释强化学习(XRL)的核心应用是调试;然而,对其相对性能的对比评估仍有限。本文提出一种新评测方法,检验用户能否从决策解释中识别智能体的目标。基于Atari的Ms. Pacman环境和四种XRL算法,结果表明:仅一种算法在测试目标上达到高于随机水平的准确率,且用户普遍对自己的判断过于自信。此外,用户自评的理解难易度与其实际识别准确率无相关性。

原文摘要 · Abstract (English)

Debugging is a core application of explainable reinforcement learning (XRL) algorithms; however, limited comparative evaluations have been conducted to understand their relative performance. We propose a novel evaluation methodology to test whether users can identify an agent's goal from an explanation of its decision-making. Utilising the Atari's Ms. Pacman environment and four XRL algorithms, we find that only one achieved greater than random accuracy for the tested goals and that users were generally overconfident in their selections. Further, we find that users' self-reported ease of identification and understanding for every explanation did not correlate with their accuracy.

XRL可解释性强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。