arXiv:2504.01482math.NAcs.LG2025-04

提出一种鲁棒方法,评估含极端事件的连续时间强化学习策略。

A Robust Model-Based Approach for Continuous-Time Policy Evaluation with Unknown Lévy Process Dynamics

  • 通过求解含未知系数的积分微分方程来评估策略价值。
  • 在真实比特币价格数据上验证了对重尾跳变过程的准确恢复能力。
  • 适合处理含极端事件的金融与复杂系统建模问题。

本文提出一种基于模型的连续时间策略评估(CTPE)框架,融合布朗运动与莱维噪声,以刻画受罕见极端事件影响的随机动态。将策略评估问题建模为求解具有未知系数的偏积分微分方程(PIDE)。核心挑战在于从轨迹数据中准确恢复未知动力学系数,尤其在莱维过程具有重尾效应时。为此,我们设计了一种鲁棒数值方法,可有效处理无偏与截断轨迹数据。该方法结合最大似然估计与迭代尾部修正机制,显著提升系数恢复的稳定性和精度。同时建立了基于系数恢复误差的策略评估误差理论界。通过数值实验,包括真实比特币价格数据实验,验证了方法在恢复重尾莱维动态方面的有效性,并确认了理论误差分析的正确性。

原文摘要 · Abstract (English)

This paper develops a model-based framework for continuous-time policy evaluation (CTPE) in reinforcement learning, incorporating both Brownian and Lévy noise to model stochastic dynamics influenced by rare and extreme events. Our approach formulates the policy evaluation problem as solving a partial integro-differential equation (PIDE) for the value function with unknown coefficients. A key challenge in this setting is accurately recovering the unknown coefficients in the stochastic dynamics, particularly when driven by Lévy processes with heavy tail effects. To address this, we propose a robust numerical approach that effectively handles both unbiased and censored trajectory datasets. This method combines maximum likelihood estimation with an iterative tail correction mechanism, improving the stability and accuracy of coefficient recovery. Additionally, we establish a theoretical bound for the policy evaluation error based on coefficient recovery error. Through numerical experiments, including a real-data BTC price experiment, we demonstrate the effectiveness and robustness of our method in recovering heavy-tailed Lévy dynamics and verify the theoretical error analysis in policy evaluation.

强化学习连续时间莱维过程金融建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。