arXiv:2506.00962cs.LGmath.OC2025-06ICML被引 1

提出随机终止时间下的强化学习新框架,提升真实场景优化效果。

Reinforcement Learning with Random Time Horizons

  • 引入轨迹依赖的随机终止时间,重新推导策略梯度公式
  • 实验显示新方法收敛速度显著优于传统算法
  • 适合动态停止条件的真实系统,如机器人任务中断

我们将标准强化学习框架扩展至随机时间终止情形。传统设定通常假设轨迹具有固定或无限时长,但许多现实应用中终止时间是随机的(可能依赖于轨迹)。由于终止时间通常依赖于策略,其随机性会影响策略梯度表达式,本文首次严谨推导了随机终止下适用于随机与确定性策略的梯度公式。我们从轨迹和状态空间两个角度提供互补视角,并建立与最优控制理论的联系。数值实验表明,使用所提公式可显著提升优化收敛性能,优于传统方法。

原文摘要 · Abstract (English)

We extend the standard reinforcement learning framework to random time horizons. While the classical setting typically assumes finite and deterministic or infinite runtimes of trajectories, we argue that multiple real-world applications naturally exhibit random (potentially trajectory-dependent) stopping times. Since those stopping times typically depend on the policy, their randomness has an effect on policy gradient formulas, which we (mostly for the first time) derive rigorously in this work both for stochastic and deterministic policies. We present two complementary perspectives, trajectory or state-space based, and establish connections to optimal control theory. Our numerical experiments demonstrate that using the proposed formulas can significantly improve optimization convergence compared to traditional approaches.

强化学习策略梯度随机过程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。