为连续动作强化学习提供可解释的反事实推理方法
Counterfactual Explanations for Continuous Action Reinforcement Learning
- 通过最小化与原动作序列偏差,生成改进结果的替代动作序列
- 在糖尿病控制和月球着陆机任务中验证了效果与泛化能力
- 适合需要高可信度决策的医疗、机器人等场景
强化学习在医疗和机器人等领域展现出巨大潜力,但因缺乏可解释性而难以应用。反事实解释能回答“如果……会怎样”的问题,是理解RL决策的有力途径,但在连续动作空间中仍研究不足。本文提出一种新颖方法,通过计算能改善结果且偏离原动作序列最小的替代动作序列,生成连续动作强化学习的反事实解释。该方法采用连续动作距离度量,并考虑特定状态下的预设策略约束。在糖尿病控制和月球着陆机两个任务中的评估表明,该方法在有效性、效率和泛化性上表现良好,推动了更可解释、更可信的强化学习应用。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) has shown great promise in domains like healthcare and robotics but often struggles with adoption due to its lack of interpretability. Counterfactual explanations, which address "what if" scenarios, provide a promising avenue for understanding RL decisions but remain underexplored for continuous action spaces. We propose a novel approach for generating counterfactual explanations in continuous action RL by computing alternative action sequences that improve outcomes while minimizing deviations from the original sequence. Our approach leverages a distance metric for continuous actions and accounts for constraints such as adhering to predefined policies in specific states. Evaluations in two RL domains, Diabetes Control and Lunar Lander, demonstrate the effectiveness, efficiency, and generalization of our approach, enabling more interpretable and trustworthy RL applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。