arXiv:2411.11697cs.LGstat.ML2024-11被引 2

针对带跳跃的连续时间强化学习,提出更鲁棒的误差评估方法。

Robust Reinforcement Learning under Diffusion Models for Data with Jumps

  • 用均方双幂变差误差替代传统TD误差,提升对跳跃过程的适应性。
  • 在带跳跃的随机微分方程环境中,价值函数估计更准确可靠。
  • 适合研究连续时间强化学习与高噪声动态系统建模的学者。

强化学习在复杂决策任务中表现优异,但在由带跳跃项的随机微分方程(SDEs)驱动的连续时间场景中仍面临挑战。本文重新审视常用于连续时间强化学习的均方TD误差(MSTDE)算法,指出其在处理状态动态跳跃时的局限性。为此,提出均方双幂变差误差(MSBVE)算法,通过最小化均方二次变差误差,显著提升在强随机噪声和跳跃过程中的鲁棒性与收敛性。仿真与形式化证明表明,该算法在带跳跃的SDE环境中能更可靠地估计价值函数,性能优于MSTDE。研究强调了采用新误差度量对提升连续时间强化学习算法韧性的重要性。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has proven effective in solving complex decision-making tasks across various domains, but challenges remain in continuous-time settings, particularly when state dynamics are governed by stochastic differential equations (SDEs) with jump components. In this paper, we address this challenge by introducing the Mean-Square Bipower Variation Error (MSBVE) algorithm, which enhances robustness and convergence in scenarios involving significant stochastic noise and jumps. We first revisit the Mean-Square TD Error (MSTDE) algorithm, commonly used in continuous-time RL, and highlight its limitations in handling jumps in state dynamics. The proposed MSBVE algorithm minimizes the mean-square quadratic variation error, offering improved performance over MSTDE in environments characterized by SDEs with jumps. Simulations and formal proofs demonstrate that the MSBVE algorithm reliably estimates the value function in complex settings, surpassing MSTDE's performance when faced with jump processes. These findings underscore the importance of alternative error metrics to improve the resilience and effectiveness of RL algorithms in continuous-time frameworks.

强化学习随机微分方程跳跃过程鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。