arXiv:2409.00328cs.LGmath.OC2024-09NeurIPS被引 21

首次实现多维奖励下可证明收敛的分布强化学习算法。

Foundations of Multivariate Distributional Reinforcement Learning

  • 提出无需先验知识的多维分布动态规划与TD学习算法。
  • 发现维度大于1时标准分析失效,引入质量为1的带符号测度投影解决。
  • 揭示不同分布表示的权衡,指导实际多目标强化学习设计。

在强化学习中,多维奖励信号的考虑推动了多目标决策、迁移学习和表征学习的根本进展。本文首次提出了无需先验知识且计算上可行的多维分布动态规划与时序差分学习的可证明收敛算法。其收敛速率与标量奖励设置下的经典速率一致,并进一步揭示了近似回报分布表示精度随奖励维度变化的特性。令人惊讶的是,当奖励维度大于1时,传统的分类TD学习分析失效,本文通过将投影映射到质量为1的带符号测度空间解决了该问题。借助技术结果与模拟实验,我们识别出影响多维分布强化学习性能的分布表示间权衡关系。

原文摘要 · Abstract (English)

In reinforcement learning (RL), the consideration of multivariate reward signals has led to fundamental advancements in multi-objective decision-making, transfer learning, and representation learning. This work introduces the first oracle-free and computationally-tractable algorithms for provably convergent multivariate distributional dynamic programming and temporal difference learning. Our convergence rates match the familiar rates in the scalar reward setting, and additionally provide new insights into the fidelity of approximate return distribution representations as a function of the reward dimension. Surprisingly, when the reward dimension is larger than $1$, we show that standard analysis of categorical TD learning fails, which we resolve with a novel projection onto the space of mass-$1$ signed measures. Finally, with the aid of our technical results and simulations, we identify tradeoffs between distribution representations that influence the performance of multivariate distributional RL in practice.

强化学习分布学习多目标算法理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。