arXiv:2512.14991cs.LGmath.OC2025-12被引 1

针对高维连续状态的随机控制问题,提出自适应分区学习算法。

Adaptive Partitioning and Learning for Stochastic Control of Diffusion Processes

  • 通过自适应划分状态-动作空间,动态优化离散化精度
  • 理论证明在高维无界环境下可实现可控后悔率
  • 适用于金融投资组合等复杂连续控制场景

我们研究具有无界连续状态空间、有界连续动作和多项式增长收益的受控扩散过程的强化学习问题,该设定常见于金融、经济与运筹领域。为应对连续高维域的挑战,提出一种基于模型的自适应分区算法,动态划分状态-动作联合空间,并在每个分区中维护漂移、波动率和收益的估计量,当估计偏差超过统计置信度时自动细化划分。该自适应机制平衡探索与近似,实现无界域内的高效学习。理论分析给出了依赖于问题时长、状态维度、收益增长阶数及新定义的‘缩放维数’的后悔上界,能覆盖有界情形的已有结果,并扩展到更广泛的扩散型问题。数值实验验证了方法有效性,包括多资产均值-方差投资组合等高维应用。

原文摘要 · Abstract (English)

We study reinforcement learning for controlled diffusion processes with unbounded continuous state spaces, bounded continuous actions, and polynomially growing rewards: settings that arise naturally in finance, economics, and operations research. To overcome the challenges of continuous and high-dimensional domains, we introduce a model-based algorithm that adaptively partitions the joint state-action space. The algorithm maintains estimators of drift, volatility, and rewards within each partition, refining the discretization whenever estimation bias exceeds statistical confidence. This adaptive scheme balances exploration and approximation, enabling efficient learning in unbounded domains. Our analysis establishes regret bounds that depend on the problem horizon, state dimension, reward growth order, and a newly defined notion of zooming dimension tailored to unbounded diffusion processes. The bounds recover existing results for bounded settings as a special case, while extending theoretical guarantees to a broader class of diffusion-type problems. Finally, we validate the effectiveness of our approach through numerical experiments, including applications to high-dimensional problems such as multi-asset mean-variance portfolio selection.

强化学习随机控制扩散过程金融建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。