arXiv:2505.12759cs.LG2025-05被引 2

用双层强化学习提升离线投资策略的泛化能力。

Your Offline Policy is Not Trustworthy: Bilevel Reinforcement Learning for Sequential Portfolio Optimization

  • 采用双层优化框架,兼顾数据内收益与跨数据变换的适应性。
  • 在两个公开股票数据集上超越现有算法,收益更优且风险更低。
  • 适合关注离线强化学习在金融场景中可靠性的研究者。

强化学习在序列化投资组合优化任务中展现巨大潜力,如股票交易,目标是利用历史数据最大化累计收益并降低风险。然而,传统强化学习方法常生成仅记忆固定数据集中最优但不切实际的买卖行为的策略,缺乏泛化能力,因其未能考虑市场的非平稳特性。本文提出MetaTrader,将投资组合优化建模为一种新型部分离线强化学习问题,并做出两项技术贡献:首先,采用双层学习框架,显式训练强化学习智能体以同时提升原始数据集内的域内收益和跨多种原始金融数据变换的域外表现;其次,引入一种新的时间差分(TD)方法,通过一批变换后的TD目标近似最坏情况下的TD估计,缓解在离线数据有限情况下常见的值函数过估计问题。在两个公开股票数据集上的实证结果表明,MetaTrader优于现有方法,包括基于强化学习的方法和传统股票预测模型。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has shown significant promise for sequential portfolio optimization tasks, such as stock trading, where the objective is to maximize cumulative returns while minimizing risks using historical data. However, traditional RL approaches often produce policies that merely memorize the optimal yet impractical buying and selling behaviors within the fixed dataset. These offline policies are less generalizable as they fail to account for the non-stationary nature of the market. Our approach, MetaTrader, frames portfolio optimization as a new type of partial-offline RL problem and makes two technical contributions. First, MetaTrader employs a bilevel learning framework that explicitly trains the RL agent to improve both in-domain profits on the original dataset and out-of-domain performance across diverse transformations of the raw financial data. Second, our approach incorporates a new temporal difference (TD) method that approximates worst-case TD estimates from a batch of transformed TD targets, addressing the value overestimation issue that is particularly challenging in scenarios with limited offline data. Our empirical results on two public stock datasets show that MetaTrader outperforms existing methods, including both RL-based approaches and traditional stock prediction models.

强化学习投资组合离线学习双层优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。