arXiv:2410.07778stat.MLcs.LG2024-10被引 5

提出网格采样SDE,用于连续时间强化学习中的探索建模。

On the grid-sampling limit SDE

  • 用网格采样构造随机微分方程模拟探索行为
  • 证明该SDE在跳跃过程存在时仍具备良好定义性
  • 为连续时间RL中的探索建模提供理论支持

在我们近期的工作[3]中,引入了网格采样SDE作为连续时间强化学习中探索行为的代理模型。本文进一步为该SDE的应用提供动机,并讨论其在跳跃过程存在情况下的适定性问题。

原文摘要 · Abstract (English)

In our recent work [3] we introduced the grid-sampling SDE as a proxy for modeling exploration in continuous-time reinforcement learning. In this note, we provide further motivation for the use of this SDE and discuss its wellposedness in the presence of jumps.

强化学习随机微分方程探索建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。