用探索式强化学习优化投机交易的买卖时机决策。
Reinforcement Learning for Speculative Trading under Exploratory Framework
- 将买卖时机建模为柯西过程的跳跃时间,通过强度控制实现随机策略。
- 引入香农微分熵正则化,得到闭式最优策略和收敛的强化学习目标。
- 适用于量化交易中的配对交易场景,适合研究随机控制与金融算法者。
我们研究了在Wang等[2020]提出的探索式强化学习框架下的投机交易问题。该问题被形式化为在一般效用函数和价格过程下,对入场与出场时间的序贯最优停止问题。首先考虑一个松弛版本,其中停止时间由受有界、非随机强度控制的柯西过程跳跃时间建模。在探索式设定下,代理的随机控制通过跳跃强度上的概率测度表征,其目标函数由香农微分熵正则化。由此导出探索式哈密顿-雅可比-贝尔曼方程,并得到闭式吉布斯最优策略。建立了误差估计及强化学习目标向原始问题价值函数的收敛性。最后设计了一个强化学习算法,并在配对交易应用中展示了其实现效果。
原文摘要 · Abstract (English)
We study a speculative trading problem within the exploratory reinforcement learning (RL) framework of Wang et al. [2020]. The problem is formulated as a sequential optimal stopping problem over entry and exit times under general utility function and price process. We first consider a relaxed version of the problem in which the stopping times are modeled by the jump times of Cox processes driven by bounded, non-randomized intensity controls. Under the exploratory formulation, the agent's randomized control is characterized via the probability measure over the jump intensities, and their objective function is regularized by Shannon's differential entropy. This yields a system of the exploratory HJB equations and Gibbs distributions in closed-form as the optimal policy. Error estimates and convergence of the RL objective to the value function of the original problem are established. Finally, an RL algorithm is designed, and its implementation is showcased in a pairs-trading application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。