arXiv:2512.02386cs.LGmath.OC2025-12被引 1

提出连续时间风险敏感强化学习算法,用于动态投资组合优化。

Risk-Sensitive Q-Learning in Continuous Time with Application to Dynamic Portfolio Selection

  • 基于随机微分方程建模连续环境,用优化确定等价物定义目标函数。
  • 证明最优策略对扩展状态为马尔可夫,且在投资组合模拟中表现有效。
  • 适合研究金融决策与风险敏感强化学习的学者和从业者。

本文研究连续时间下的风险敏感强化学习(RSRL)问题,其中环境由可控随机微分方程(SDE)刻画,目标函数为累积回报的非线性泛函。当该泛函为优化确定等价物(OCE)时,我们证明最优策略关于扩展状态是马尔可夫的。为此,我们提出一种名为CT-RS-q的风险敏感Q-learning算法,基于新颖的鞅表征方法。最后,在动态投资组合选择问题上进行仿真研究,验证了算法的有效性。

原文摘要 · Abstract (English)

This paper studies the problem of risk-sensitive reinforcement learning (RSRL) in continuous time, where the environment is characterized by a controllable stochastic differential equation (SDE) and the objective is a potentially nonlinear functional of cumulative rewards. We prove that when the functional is an optimized certainty equivalent (OCE), the optimal policy is Markovian with respect to an augmented environment. We also propose \textit{CT-RS-q}, a risk-sensitive q-learning algorithm based on a novel martingale characterization approach. Finally, we run a simulation study on a dynamic portfolio selection problem and illustrate the effectiveness of our algorithm.

强化学习金融建模连续时间风险敏感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。