arXiv:2607.16999cs.LGcs.AI2026-07

用反事实方法精准区分智能体技能与环境运气,提升强化学习可解释性。

Counterfactual Shapley Credit Assignment

论文配图:Counterfactual Shapley Credit Assignment
图 1 · 摘自论文原文
  • 基于因果理论设计反事实谢尔利值,分离技能与运气的影响。
  • 在稀疏奖励、高随机性环境下,样本效率显著优于现有方法。
  • 适合需要高可解释性的强化学习系统研发者使用。

信用分配问题(CAP)是构建高效且可解释的强化学习智能体的核心挑战。现有方法依赖时间连续性或事后奖励重加权,常无法正确区分智能体策略(技能)与环境随机性(运气)。我们提出反事实谢尔利信用分配框架,基于因果理论,通过反事实谢尔利值(ϕ-值)实现信用与责备的精准分配。该方法通过重新分配环境奖励,在稀疏因果、高随机性和延迟奖励三个关键维度上显著改善时间信用分配,同时保持最优策略不变。我们推导出一个一致的估计器,高效计算ϕ-值,进而构建新型策略梯度方法ϕ-PPO,结合优先级轨迹重放(PTR)。实验表明,ϕ-值能精确匹配任务奖励的真实成因,在以往最先进方法无法收敛的复杂环境中展现出更优的样本效率。

原文摘要 · Abstract (English)

The Credit Assignment Problem (CAP) is fundamental to developing efficient and explainable Reinforcement Learning (RL) agents. Existing frameworks, whether relying on temporal contiguity or hindsight-conditioned reward reweighting, frequently fail to attribute properly between an agent's policy (skill) and environmental stochasticity (luck). A principled approach to CAP must isolate the true causal drivers of observed outcomes from spurious correlations and environmental randomness. We introduce Counterfactual Shapley Credit Assignment, a novel framework grounded in causal theory that attributes credit and blame via the Counterfactual Shapley Value ($ϕ$-value). By redistributing environmental rewards, $ϕ$-values enhance temporal credit assignment across three critical dimensions: sparse causality, high stochasticity, and delayed rewards, all while preserving the optimal policy. We derive a consistent estimator that computes $ϕ$-values efficiently, enabling a new class of policy gradient methods, $ϕ$-PPO, combined with Prioritized Trajectory Replay (PTR). Empirical results demonstrate that $ϕ$-values align precisely to the ground truth causes of task rewards with superior sample efficiency in challenging environments where prior state-of-the-art methods fail to converge.

强化学习信用分配因果推理可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。