arXiv:2512.06917cs.LG2025-12中稿 · AAAI

通过轨迹重要性分析,让强化学习决策更可解释、可信。

Know your Trajectory -- Trustworthy Reinforcement Learning deployment through Importance-Based Trajectory Analysis

  • 用新定义的状态重要性度量,综合评估状态对目标达成的影响。
  • 在标准环境上,比传统方法更准确识别最优轨迹。
  • 适合需要透明决策的自动驾驶、机器人等安全敏感场景。

随着强化学习(RL)在真实场景中广泛应用,确保其行为透明可信至关重要。现有可解释强化学习(XRL)多聚焦单步决策,难以解释长期行为。本文提出一种新框架,通过定义并聚合新的状态重要性度量,对完整轨迹进行排序。该度量结合经典Q值差与一个捕捉代理向目标靠近倾向的‘激进项’,提供更细致的状态关键性评估。实验表明,该方法能从异构经验中成功识别最优轨迹。进一步通过关键状态生成反事实回放,证明所选路径在多种替代方案中均具鲁棒优势,实现‘为何选此而非彼’的有力解释。在标准OpenAI Gym环境中验证,该重要性度量显著优于传统方法,推动可信自主系统发展。

原文摘要 · Abstract (English)

As Reinforcement Learning (RL) agents are increasingly deployed in real-world applications, ensuring their behavior is transparent and trustworthy is paramount. A key component of trust is explainability, yet much of the work in Explainable RL (XRL) focuses on local, single-step decisions. This paper addresses the critical need for explaining an agent's long-term behavior through trajectory-level analysis. We introduce a novel framework that ranks entire trajectories by defining and aggregating a new state-importance metric. This metric combines the classic Q-value difference with a "radical term" that captures the agent's affinity to reach its goal, providing a more nuanced measure of state criticality. We demonstrate that our method successfully identifies optimal trajectories from a heterogeneous collection of agent experiences. Furthermore, by generating counterfactual rollouts from critical states within these trajectories, we show that the agent's chosen path is robustly superior to alternatives, thereby providing a powerful "Why this, and not that?" explanation. Our experiments in standard OpenAI Gym environments validate that our proposed importance metric is more effective at identifying optimal behaviors compared to classic approaches, offering a significant step towards trustworthy autonomous systems.

强化学习可解释性轨迹分析可信部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。