arXiv:2602.17098q-fin.PMcs.AI2026-02被引 24

对比深度强化学习与传统均值-方差法,验证DRL在投资组合优化中的优越性。

Deep Reinforcement Learning for Optimal Portfolio Allocation: A Comparative Study with Mean-Variance Optimization

  • 采用无模型强化学习训练代理,基于历史数据自动学习最优资产配置策略。
  • DRL在夏普比率、最大回撤和绝对收益上均显著优于均值-方差法。
  • 为金融从业者提供可落地的DRL实操框架,适合量化投资研究者参考。

投资组合管理是为实现预定投资目标而统筹管理一组资产的过程。投资组合优化是其中关键环节,旨在通过合理分配资产以最大化收益并最小化风险,通常由金融专业人士结合定量方法与经验决策完成。近年来,深度强化学习(DRL)在利用历史市场数据训练无模型智能体方面展现出良好前景。然而,多数方法仅与基础基准或先进DRL模型比较,却很少与金融实践中常用的均值-方差优化(MVO)进行对照。MVO通过历史时间序列估计资产预期收益与协方差,用于优化投资目标。本文系统比较了无模型DRL与MVO在投资组合优化中的表现。详细说明了DRL实际应用的具体实现方式,并指出MVO需进行的调整。回测结果表明,DRL代理在夏普比率、最大回撤和绝对收益等多个指标上均表现优异。

原文摘要 · Abstract (English)

Portfolio Management is the process of overseeing a group of investments, referred to as a portfolio, with the objective of achieving predetermined investment goals. Portfolio optimization is a key component that involves allocating the portfolio assets so as to maximize returns while minimizing risk taken. It is typically carried out by financial professionals who use a combination of quantitative techniques and investment expertise to make decisions about the portfolio allocation. Recent applications of Deep Reinforcement Learning (DRL) have shown promising results when used to optimize portfolio allocation by training model-free agents on historical market data. Many of these methods compare their results against basic benchmarks or other state-of-the-art DRL agents but often fail to compare their performance against traditional methods used by financial professionals in practical settings. One of the most commonly used methods for this task is Mean-Variance Portfolio Optimization (MVO), which uses historical time series information to estimate expected asset returns and covariances, which are then used to optimize for an investment objective. Our work is a thorough comparison between model-free DRL and MVO for optimal portfolio allocation. We detail the specifics of how to make DRL for portfolio optimization work in practice, also noting the adjustments needed for MVO. Backtest results demonstrate strong performance of the DRL agent across many metrics, including Sharpe ratio, maximum drawdowns, and absolute returns.

强化学习投资组合量化交易DRL

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。