arXiv:2608.20680cs.LG2026-08

提出适用于网络动态定价的连续时间强化学习算法,解决离散状态空间下的策略优化难题。

Reinforcement Learning for Continuous-Time Jump Markov Decision Processes with Applications to Network Dynamic Pricing

  • 基于熵正则化构建连续时间控制模型,适配离散状态空间的强化学习框架。
  • 在真实网络动态定价场景中,算法学习到接近最优策略,性能显著优于传统方法。
  • 适用于资源受限的多产品定价场景,适合大规模系统应用者参考。

本文研究具有通用离散状态空间(无需向量空间结构)和连续/离散动作空间的连续时间跳跃马尔可夫决策过程(CTJMDP)中的强化学习问题。该框架涵盖运营领域诸多经典应用,如资源容量受限的多产品动态定价(Gallego and van Ryzin, 1997)。为建模探索与利用的权衡,我们提出带有随机策略的熵正则化连续时间控制问题。现有连续时间强化学习方法(如Jia and Zhou 2023提出的扩散过程q-learning)主要针对ℝ^d连续状态空间,依赖ℝ^d中的半鞅理论进行分析,难以直接应用于一般离散状态空间的CTJMDP,因后者可能缺乏欧几里得空间固有的加减结构。为此,本文建立了CTJMDP下q-learning的理论基础,并开发了无模型q-learning算法。相比朴素的时间离散化及将CTJMDP近似为离散时间MDP的方法,本方法在概念和实证层面均有优势。在网络动态定价实验(Gallego and van Ryzin, 1997)中,所提算法能稳定学习近优策略,持续优于标准基准方法,展现出更优解质量与对大规模网络实例的有效扩展性。

原文摘要 · Abstract (English)

We study reinforcement learning (RL) in Continuous-Time Jump Markov Decision Processes (CTJMDPs) featuring general discrete state spaces (which need not possess a vector space structure) and continuous/discrete action spaces. The setup covers many well-known applications in operations such as multi-product dynamic pricing with capacitated resources (Gallego and van Ryzin 1997). To model the exploration-exploitation tradeoff, we formulate an entropy-regularized continuous-time control problem with stochastic policies. Recent continuous-time RL techniques such as $q$-learning for controlled diffusions in (Jia and Zhou 2023) focus on continuous state spaces $\mathbb{R}^d$ and rely heavily on semimartingale theory in $\mathbb{R}^d$ for their theoretical analysis. Consequently, their methods cannot be directly applied to CTJMDPs with general discrete state spaces, which may lack the algebraic addition and subtraction structures inherent to Euclidean spaces. To bridge this gap, we establish the theoretical foundations of $q$-learning for CTJMDPs and develop model-free $q$-learning algorithms. Compared to naïve time discretization and approximating CTJMDPs using discrete-time MDPs, our approach has several conceptual and empirical benefits. Numerical experiments in network dynamic pricing (Gallego and van Ryzin 1997) show that our proposed RL algorithm reliably learns near-optimal policies and consistently outperforms standard benchmark methods, demonstrating superior solution quality and effective scalability to large-scale network instances.

强化学习动态定价连续时间网络优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。