用模仿学习提升金融交易中噪声环境下的强化学习表现
Robot See, Robot Do: Imitation Reward for Noisy Financial Environments
- 引入趋势标注算法作为专家,融合模仿与强化反馈
- 在真实市场数据上,收益与夏普比率显著优于传统方法
- 适合处理高噪声金融环境的策略优化问题
金融资产交易的序列决策特性天然契合强化学习框架,但市场低信噪比导致奖励函数等环境组件估计噪声大,阻碍了强化学习代理的有效策略学习。鉴于奖励函数设计在强化学习中的关键作用,本文提出一种新方法:利用模仿学习构建更鲁棒的奖励函数,其中趋势标注算法充当专家。将模仿(专家)反馈与强化(代理)反馈融合进一个无模型强化学习算法中,有效将模仿学习问题嵌入强化学习范式,以应对奖励信号的随机性。实验结果表明,该方法在财务绩效指标上优于传统基准及仅使用强化反馈训练的强化学习代理。
原文摘要 · Abstract (English)
The sequential nature of decision-making in financial asset trading aligns naturally with the reinforcement learning (RL) framework, making RL a common approach in this domain. However, the low signal-to-noise ratio in financial markets results in noisy estimates of environment components, including the reward function, which hinders effective policy learning by RL agents. Given the critical importance of reward function design in RL problems, this paper introduces a novel and more robust reward function by leveraging imitation learning, where a trend labeling algorithm acts as an expert. We integrate imitation (expert's) feedback with reinforcement (agent's) feedback in a model-free RL algorithm, effectively embedding the imitation learning problem within the RL paradigm to handle the stochasticity of reward signals. Empirical results demonstrate that this novel approach improves financial performance metrics compared to traditional benchmarks and RL agents trained solely using reinforcement feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。