arXiv:2504.14645cs.LGcs.AI2025-04被引 2

用代理评估指标优化强化学习示范轨迹,提升策略可解释性。

Surrogate Fitness Metrics for Interpretable Reinforcement Learning

  • 设计联合代理适应度函数,融合局部多样性、行为确定性和全局种群多样性。
  • 在网格世界和连续控制任务中,示范保真度显著优于随机与消融基线。
  • 适合关注安全关键与可解释性场景的RL研究者参考。

我们采用进化优化框架,通过扰动初始状态生成信息丰富且多样化的策略示范。联合代理适应度函数结合局部多样性、行为确定性和全局种群多样性来引导优化。为评估示范质量,引入奖励最优性差距、保真度四分位均值(IQMs)、适应度构成分析及轨迹可视化等指标。同时考察超参数敏感性以理解轨迹优化动态。实验表明,在离散与连续环境中,基于代理适应度指标优化轨迹选择显著提升强化学习策略的可解释性。在网格世界中,示范保真度显著优于随机与消融基线;在连续控制任务中,该框架对早期策略提供有效洞察,而保真度优化对成熟策略更优。通过系统改进与分析代理适应度函数,本研究推动了强化学习模型可解释性的进展,深化了对决策过程的理解,适用于安全关键与可解释性导向的应用场景。

原文摘要 · Abstract (English)

We employ an evolutionary optimization framework that perturbs initial states to generate informative and diverse policy demonstrations. A joint surrogate fitness function guides the optimization by combining local diversity, behavioral certainty, and global population diversity. To assess demonstration quality, we apply a set of evaluation metrics, including the reward-based optimality gap, fidelity interquartile means (IQMs), fitness composition analysis, and trajectory visualizations. Hyperparameter sensitivity is also examined to better understand the dynamics of trajectory optimization. Our findings demonstrate that optimizing trajectory selection via surrogate fitness metrics significantly improves interpretability of RL policies in both discrete and continuous environments. In gridworld domains, evaluations reveal significantly enhanced demonstration fidelities compared to random and ablated baselines. In continuous control, the proposed framework offers valuable insights, particularly for early-stage policies, while fidelity-based optimization proves more effective for mature policies. By refining and systematically analyzing surrogate fitness functions, this study advances the interpretability of RL models. The proposed improvements provide deeper insights into RL decision-making, benefiting applications in safety-critical and explainability-focused domains.

强化学习可解释性进化优化代理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。