强化学习能发现价格操纵机会,且在参数不准时比传统方法更有效。
Can Reinforcement Learning Efficiently Discover Price Manipulation?

- 用深度确定性策略梯度直接学操纵策略,不依赖模型假设。
- 中等波动下,有限数据中强化学习仍能发现盈利策略。
- 适合研究金融算法风险或探索无模型控制的学者。
本文研究无模型强化学习(RL)代理能否比基于模型的方法更有效地发现价格操纵机会。考虑单资产市场,价格遵循带有非线性永久影响和线性临时影响的Almgren-Chriss框架。首先在离散时间下证明了操纵策略的存在性,并通过序贯最小二乘二次规划(Sequential Least Squares Quadratic Programming)计算出理想基准策略。随后比较两种有限样本学习方法:基于模拟执行数据估计影响参数的模型方法,以及直接使用相同数据训练的深度确定性策略梯度(Deep Deterministic Policy Gradient)的无模型RL方法。在中等波动下,即使训练数据有限,RL代理也能在未明确知晓底层模型的情况下成功发现盈利操纵策略。更重要的是,当参数估计受抽样误差影响时,尽管模型方法具备正确模型设定优势,但RL始终表现更优。在高波动下,所有方法均无法识别操纵机会;低波动下,模型方法优于RL。结果凸显了强化学习在复杂控制问题中的有效性,也揭示了在金融市场部署学习算法时缺乏适当保障的风险。
原文摘要 · Abstract (English)
In this paper, we investigate whether a model-free RL agent can identify and exploit price manipulation opportunities more effectively than a traditional model-based approach that assumes correct specification of the data-generating process but relies on noisy parameter estimates. We consider a single-asset market in which prices evolve according to an Almgren-Chriss framework with non-linear permanent impact and linear temporary impact. We first establish the existence of price-manipulative strategies in discrete time and compute the optimal benchmark strategy using Sequential Least Squares Quadratic Programming under full information. We then compare two finite-sample learning approaches: a model-based procedure that estimates impact parameters from simulated execution data and an agnostic RL approach based on Deep Deterministic Policy Gradient, trained directly on the same amount of data. For intermediate volatility, the RL agent successfully discovers profitable manipulative strategies without explicit knowledge of the underlying model, even when training data are quite limited. More importantly, RL consistently outperforms the model-based approach when parameter estimates are affected by sampling error, despite the latter benefiting from the correct model specification. For large volatility, all methods are unable to identify manipulation opportunities, while for small volatility, the model based approach outperforms RL. These findings highlight both the effectiveness of RL in complex control problems and the risks associated with deploying learning algorithms in financial markets without appropriate safeguards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。