arXiv:2506.09901cs.LG2025-06

让强化学习模型给出多种合理可选的行动路径,帮助人类理解决策。

"What are my options?": Explaining RL Agents with Diverse Near-Optimal Alternatives (Extended)

  • 通过局部奖励重塑,生成在欧氏空间中差异明显的多条最优策略。
  • 所获策略均保证ε-最优性,且在模拟中表现出显著不同的轨迹形状。
  • 适合需要解释性与灵活规划的RL应用场景,如机器人路径选择。

本文扩展讨论了一种名为多样近优选项(Diverse Near-Optimal Alternatives, DNA)的可解释强化学习新方法,首次提出于L4DC 2025。DNA旨在为轨迹规划类代理生成一组合理的‘选项’,通过优化策略以在欧氏空间中产生定性上多样的轨迹。基于可解释性理念,这些不同策略可帮助人类用户从多种可用轨迹形态中进行选择。该方法适用于基于价值函数的马尔可夫决策过程中的连续轨迹代理。本文描述了DNA机制,其通过在局部修改的Q-learning问题中使用奖励塑造,求解具有保证ε-最优性的不同策略。实验表明,该方法能成功返回在模拟中具有明显区别的策略,构成有意义的‘选项’,并简要对比了与质量多样性(Quality Diversity)领域相关方法的异同。除解释动机外,该工作也为强化学习中的探索与自适应规划开辟了新可能。

原文摘要 · Abstract (English)

In this work, we provide an extended discussion of a new approach to explainable Reinforcement Learning called Diverse Near-Optimal Alternatives (DNA), first proposed at L4DC 2025. DNA seeks a set of reasonable "options" for trajectory-planning agents, optimizing policies to produce qualitatively diverse trajectories in Euclidean space. In the spirit of explainability, these distinct policies are used to "explain" an agent's options in terms of available trajectory shapes from which a human user may choose. In particular, DNA applies to value function-based policies on Markov decision processes where agents are limited to continuous trajectories. Here, we describe DNA, which uses reward shaping in local, modified Q-learning problems to solve for distinct policies with guaranteed epsilon-optimality. We show that it successfully returns qualitatively different policies that constitute meaningfully different "options" in simulation, including a brief comparison to related approaches in the stochastic optimization field of Quality Diversity. Beyond the explanatory motivation, this work opens new possibilities for exploration and adaptive planning in RL.

强化学习可解释性路径规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。