arXiv:2503.08322cs.LGcs.AI2025-03被引 3

用程序化策略蒸馏法,无需真人测试就能评估强化学习的可解释性。

Evaluating Interpretable Reinforcement Learning by Distilling Policies into Programs

  • 将专家神经网络蒸馏为小型可读程序,作为可解释策略基线。
  • 方法评估结果与真人研究一致,且可解释性不降低甚至提升性能。
  • 揭示无通用最优可解释性策略,适合需权衡可解释与性能的研究者。

在医疗等应用中,强化学习策略需具备人类可理解性。已有用户研究表明某些策略类型更易解释,但人工评估成本高且缺乏统一定义。本文提出一种无需人类参与的可解释性评估新方法,基于「可模拟性」这一共识概念——即人类能否根据状态理解策略行为。通过模仿学习,将专家神经网络蒸馏为小型程序作为基线策略,并在此基础上开展大规模实证评估。结果表明,该方法得出的结论与用户研究一致;提升可解释性并不必然导致性能下降,甚至可能提升;且不存在在所有任务上均优的策略类别,因此需要可比较的评估工具。本方法推动了可解释强化学习研究的发展。

原文摘要 · Abstract (English)

There exist applications of reinforcement learning like medicine where policies need to be ''interpretable'' by humans. User studies have shown that some policy classes might be more interpretable than others. However, it is costly to conduct human studies of policy interpretability. Furthermore, there is no clear definition of policy interpretabiliy, i.e., no clear metrics for interpretability and thus claims depend on the chosen definition. We tackle the problem of empirically evaluating policies interpretability without humans. Despite this lack of clear definition, researchers agree on the notions of ''simulatability'': policy interpretability should relate to how humans understand policy actions given states. To advance research in interpretable reinforcement learning, we contribute a new methodology to evaluate policy interpretability. This new methodology relies on proxies for simulatability that we use to conduct a large-scale empirical evaluation of policy interpretability. We use imitation learning to compute baseline policies by distilling expert neural networks into small programs. We then show that using our methodology to evaluate the baselines interpretability leads to similar conclusions as user studies. We show that increasing interpretability does not necessarily reduce performances and can sometimes increase them. We also show that there is no policy class that better trades off interpretability and performance across tasks making it necessary for researcher to have methodologies for comparing policies interpretability.

可解释RL策略蒸馏评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。