提出PAMC方法,用低秩+稀疏结构提升稀疏奖励学习效率
What Fundamental Structure in Reward Functions Enables Efficient Sparse-Reward Learning?
- 利用策略偏倚采样下的低秩加稀疏结构建模奖励矩阵
- 在多个基准上实现1000倍样本效率提升,超越主流算法
- 支持安全退化,适合有结构奖励的实际强化学习场景
稀疏奖励强化学习本质困难:无结构时,任何智能体需Ω(|S||A|/p)次采样才能恢复奖励。我们提出策略感知矩阵补全(PAMC),作为结构化奖励学习框架的首个具体实例。核心思想是在策略偏倚(MNAR)采样下,利用奖励矩阵的近似低秩加稀疏结构。通过反倾向加权证明了恢复保证,并建立访问加权误差-遗憾边界,将补全误差与控制性能关联。重要的是,当假设弱化时,PAMC可渐进退化:置信区间扩大且算法放弃预测,确保探索安全。实验表明,PAMC在Atari-26(10M步)、DM Control、MetaWorld MT50、D4RL离线强化学习及基于偏好强化学习基准中均显著提升样本效率,优于DrQ-v2、DreamerV3、Agent57、T-REX/D-REX和PrefPPO,在计算量归一化比较下表现更优。结果表明,当存在结构化奖励时,PAMC是一种实用且原理严谨的工具,也是更广泛结构化奖励学习视角的首个具体实现。
原文摘要 · Abstract (English)
Sparse-reward reinforcement learning (RL) remains fundamentally hard: without structure, any agent needs $Ω(|\mathcal{S}||\mathcal{A}|/p)$ samples to recover rewards. We introduce Policy-Aware Matrix Completion (PAMC) as a first concrete step toward a structural reward learning framework. Our key idea is to exploit approximate low-rank + sparse structure in the reward matrix, under policy-biased (MNAR) sampling. We prove recovery guarantees with inverse-propensity weighting, and establish a visitation-weighted error-to-regret bound linking completion error to control performance. Importantly, when assumptions weaken, PAMC degrades gracefully: confidence intervals widen and the algorithm abstains, ensuring safe fallback to exploration. Empirically, PAMC improves sample efficiency across Atari-26 (10M steps), DM Control, MetaWorld MT50, D4RL offline RL, and preference-based RL benchmarks, outperforming DrQ-v2, DreamerV3, Agent57, T-REX/D-REX, and PrefPPO under compute-normalized comparisons. Our results highlight PAMC as a practical and principled tool when structural rewards exist, and as a concrete first instantiation of a broader structural reward learning perspective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。