arXiv:2601.02514cs.AI2026-01

让强化学习决策可解释,生成可验证的规则化文本说明。

Textual Explanations and Their Evaluations for Reinforcement Learning Policy

  • 用大模型与聚类生成常见状态条件,转为透明规则
  • 提出两种优化技术,提升解释一致性和准确性
  • 在3个开源环境和电信场景验证,支持量化评估

理解强化学习策略对确保自主代理行为符合人类预期至关重要。本文提出一种新型可解释强化学习(XRL)框架,生成易于理解的文本解释,并将其转化为透明规则以提升质量与可验证性。通过引入专家知识与自动谓词生成器,识别状态的语义信息;利用大语言模型与聚类技术提取频繁状态条件,进而生成规则并评估其属性、保真度与部署表现。提出两种精炼技术以减少冲突信息。在三个开源环境与一个电信工业场景中进行实验,验证了框架的可复现性与实用性。相比现有方法(如Autonomous Policy Explanation),该框架在特定任务上实现良好性能,并提供系统化、量化的文本解释评估方式,推动XRL领域发展。

原文摘要 · Abstract (English)

Understanding a Reinforcement Learning (RL) policy is crucial for ensuring that autonomous agents behave according to human expectations. This goal can be achieved using Explainable Reinforcement Learning (XRL) techniques. Although textual explanations are easily understood by humans, ensuring their correctness remains a challenge, and evaluations in state-of-the-art remain limited. We present a novel XRL framework for generating textual explanations, converting them into a set of transparent rules, improving their quality, and evaluating them. Expert's knowledge can be incorporated into this framework, and an automatic predicate generator is also proposed to determine the semantic information of a state. Textual explanations are generated using a Large Language Model (LLM) and a clustering technique to identify frequent conditions. These conditions are then converted into rules to evaluate their properties, fidelity, and performance in the deployed environment. Two refinement techniques are proposed to improve the quality of explanations and reduce conflicting information. Experiments were conducted in three open-source environments to enable reproducibility, and in a telecom use case to evaluate the industrial applicability of the proposed XRL framework. This framework addresses the limitations of an existing method, Autonomous Policy Explanation, and the generated transparent rules can achieve satisfactory performance on certain tasks. This framework also enables a systematic and quantitative evaluation of textual explanations, providing valuable insights for the XRL field.

可解释AI强化学习文本生成规则提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。