自动发现辅助任务,让强化学习炒股更稳定
Self-Supervised Auxiliary Task Discovery for Stable Reinforcement Learning in Stock Trading
- 用自监督方法自动生成有助于学习的辅助任务
- 在4个股指上验证,策略稳定性与收益均优于基线
- 适合关注强化学习交易鲁棒性的研究者
强化学习在股票交易中日益受到关注,但因市场非平稳性和奖励信号噪声,实现盈利且稳定的策略仍具挑战。传统辅助任务多为人工设计,依赖先验假设,难以适应市场变化。本文提出一种自监督框架,自动发现用于支持强化学习的辅助任务。这些任务被建模为广义价值函数,其预测可丰富状态表示并辅助策略优化。框架包含两个网络:主网络学习交易策略及辅助预测,辅网络通过学习到的累积奖赏和折扣因子生成任务定义。任务更新采用元梯度机制,考虑其对长期交易表现的影响,提升训练稳定性。在道琼斯、富时、孟买敏感指数和台湾加权指数四个主要股指上进行评估,结果表明,自动发现的辅助任务显著增强学习鲁棒性并提升交易绩效。
原文摘要 · Abstract (English)
Reinforcement learning has gained increasing attention as a data-driven approach for stock trading. However, learning a policy that is both profitable and stable remains challenging due to non-stationary market behaviour and noisy reward signals. Auxiliary tasks are often used to improve representation learning and stabilize training, yet they are usually designed manually and depend heavily on prior assumptions about targets and prediction horizons. Such fixed designs may not remain suitable across changing market regimes. In this work, we propose a self-supervised framework that automatically discovers auxiliary tasks to support reinforcement learning for stock trading. The auxiliary tasks are formulated as General Value Functions so that their predictions enrich the learned state representation and assist policy optimization. The framework consists of two networks. The main network learns the trading policy along with the auxiliary predictions, while the secondary network generates the definitions of auxiliary tasks through learned cumulants and discount factors. These tasks are updated using a meta gradient mechanism that accounts for their long-term impact on trading performance and improves training stability. We evaluate the proposed approach across four major equity indices: DJI, FTSE, Sensex, and TAIEX. The empirical results demonstrate that automatically discovered auxiliary tasks lead to more robust learning and improved trading performance compared to existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。