arXiv:2604.13147stat.MLcs.LG2026-04

针对非马尔可夫随机控制问题,提出无需重采样的自适应学习方法。

Adaptive Learning via Off-Model Training and Importance Sampling for Fully Non-Markovian Optimal Stochastic Control. Complete version

论文配图:Adaptive Learning via Off-Model Training and Importance Sampling for Fully Non-Markovian Optimal Stochastic Control. Complete version
图 1 · 摘自论文原文
  • 通过重要性采样在固定数据集上训练,实现离模学习
  • 在线性二次模型中验证了误差可控的自适应更新机制
  • 适合处理参数不确定的金融衍生品对冲等场景

本文研究连续时间随机控制问题,其受控状态为完全非马尔可夫且依赖未知模型参数。此类问题自然出现在路径相关随机微分方程、粗糙波动率对冲及分数布朗运动驱动系统中。基于前期工作的离散骨架方法,我们提出了求解嵌入式后向动态规划方程的蒙特卡洛学习方法。主要贡献有二:第一,为几类典型非马尔可夫受控系统构造显式主导训练律与Radon-Nikodym权重,实现固定合成数据集下的离模训练;第二,利用该结构设计参数不确定性下的自适应更新机制,通过重加权现有样本实现重复校准,无需重新生成轨迹。对于固定参数,建立了深度神经网络近似嵌入式动态规划方程的非渐近误差界;对于自适应学习,推导出蒙特卡洛逼近误差与模型风险误差的定量分离估计。数值实验展示了离模训练机制与自适应重要性采样更新在结构化线性二次模型中的有效性。

原文摘要 · Abstract (English)

This paper studies continuous-time stochastic control problems whose controlled states are fully non-Markovian and depend on unknown model parameters. Such problems arise naturally in path-dependent stochastic differential equations, rough-volatility hedging, and systems driven by fractional Brownian motion. Building on the discrete skeleton approach developed in earlier work, we propose a Monte Carlo learning methodology for the associated embedded backward dynamic programming equation. Our main contribution is twofold. First, we construct explicit dominating training laws and Radon--Nikodym weights for several representative classes of non-Markovian controlled systems. This yields an off-model training architecture in which a fixed synthetic dataset is generated under a reference law, while the dynamic programming operators associated with a target model are recovered by importance sampling. Second, we use this structure to design an adaptive update mechanism under parametric model uncertainty, so that repeated recalibration can be performed by reweighting the same training sample rather than regenerating new trajectories. For fixed parameters, we establish non-asymptotic error bounds for the approximation of the embedded dynamic programming equation via deep neural networks. For adaptive learning, we derive quantitative estimates that separate Monte Carlo approximation error from model-risk error. Numerical experiments illustrate both the off-model training mechanism and the adaptive importance-sampling update in structured linear-quadratic examples.

随机控制非马尔可夫重要性采样自适应学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。