arXiv:2504.05625cs.LG2025-04被引 5

无需模型内部信息,用大模型生成可理解的智能体行为解释。

Model-Agnostic Policy Explanations with Large Language Models

  • 通过观察状态和动作构建局部可解释代理模型,指导大模型生成解释。
  • 生成的解释在可懂性和正确性上优于基线方法,人类和模型评估均认可。
  • 用户研究显示,该解释能提升对智能体未来行为的预测准确率。

智能体(如机器人)越来越多地部署在以人类为中心的真实环境中。为建立人类信任并满足法律与伦理要求,这些智能体必须能够解释自身行为。然而,当前最先进的智能体通常由深度神经网络等黑箱模型驱动,限制了其可解释性。本文提出一种仅基于观测到的状态和动作生成自然语言解释的方法——无需访问智能体的底层模型。该方法从观测数据中学习一个局部可解释的代理模型,再以此引导大语言模型生成合理且幻觉极少的解释。实验结果表明,相比基线方法,该方法生成的解释在语言模型和人工评估中均更易懂、更准确。此外,用户研究表明,获得本文解释的参与者能更准确预测智能体未来的行动,说明其显著提升了对智能体行为的理解。

原文摘要 · Abstract (English)

Intelligent agents, such as robots, are increasingly deployed in real-world, human-centric environments. To foster appropriate human trust and meet legal and ethical standards, these agents must be able to explain their behavior. However, state-of-the-art agents are typically driven by black-box models like deep neural networks, limiting their interpretability. We propose a method for generating natural language explanations of agent behavior based only on observed states and actions -- without access to the agent's underlying model. Our approach learns a locally interpretable surrogate model of the agent's behavior from observations, which then guides a large language model to generate plausible explanations with minimal hallucination. Empirical results show that our method produces explanations that are more comprehensible and correct than those from baselines, as judged by both language models and human evaluators. Furthermore, we find that participants in a user study more accurately predicted the agent's future actions when given our explanations, suggesting improved understanding of agent behavior.

可解释AI大模型应用行为解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。