arXiv:2505.08073cs.AI2025-05中稿 · Workshop on Explai…被引 2

用世界模型生成可解释的强化学习决策原因,提升用户理解力。

Explainable Reinforcement Learning Agents Using World Models

  • 构建反向世界模型,逆推理想状态以解释为何选择某动作
  • 通过反事实轨迹对比,使用户理解策略行为背后的逻辑
  • 适合非专家用户学习如何通过环境干预控制智能体

可解释人工智能(XAI)旨在帮助人们理解AI系统的输出与行为。可解释强化学习(XRL)因序列决策的时间特性而更具挑战性,且非专业用户通常无法修改智能体或其策略。本文提出利用世界模型为基于模型的深度强化学习智能体生成解释。世界模型可预测执行动作后世界的变化,从而生成反事实轨迹。然而,仅知用户期望的行为仍不足以理解智能体为何做出不同选择。为此,我们为基于模型的强化学习智能体引入反向世界模型,该模型可预测在何种理想状态下,智能体才会偏好某个特定的反事实动作。实验表明,展示‘世界本应如何’的解释显著提升了用户对智能体策略的理解。我们假设此类解释能帮助用户通过操纵环境来学会控制智能体的执行过程。

原文摘要 · Abstract (English)

Explainable AI (XAI) systems have been proposed to help people understand how AI systems produce outputs and behaviors. Explainable Reinforcement Learning (XRL) has an added complexity due to the temporal nature of sequential decision-making. Further, non-AI experts do not necessarily have the ability to alter an agent or its policy. We introduce a technique for using World Models to generate explanations for Model-Based Deep RL agents. World Models predict how the world will change when actions are performed, allowing for the generation of counterfactual trajectories. However, identifying what a user wanted the agent to do is not enough to understand why the agent did something else. We augment Model-Based RL agents with a Reverse World Model, which predicts what the state of the world should have been for the agent to prefer a given counterfactual action. We show that explanations that show users what the world should have been like significantly increase their understanding of the agent policy. We hypothesize that our explanations can help users learn how to control the agents execution through by manipulating the environment.

可解释AI强化学习世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。