arXiv:2509.16206cs.CEcs.LG2025-09

用深度强化学习构建因子投资组合,显著提升长期收益与风险收益比。

Deep Reinforcement Learning in Factor Investment

  • 将个股收益压缩为94个特征条件下的低维因子,缓解高维状态空间问题
  • 在20年美股数据上实现24.6%年化复合收益和0.94的夏普比率
  • 方法适合机构级低换手率投资,且可解释性良好

深度强化学习在交易执行中已显成效,但在低频因子投资组合构建中仍少有应用。主要瓶颈在于股票进出投资组合导致的高维、不平衡状态空间。本文提出条件自编码因子型组合优化(CAFPO),通过94个公司特有特征对个股收益进行压缩,生成少量潜在因子,并输入基于PPO和DDPG的DRL代理以生成连续多空权重。在2000至2020年的美国股市数据上,CAFPO超越等权、市值加权、马科维茨、原始DRL及法玛-弗伦奇驱动的DRL,实现24.6%的年化复合收益率和0.94的样本外夏普比率。SHAP分析揭示了具有经济意义的因子贡献。结果表明,因子感知的表征学习可使DRL适用于机构级低换手率组合管理。

原文摘要 · Abstract (English)

Deep reinforcement learning has shown promise in trade execution, yet its use in low-frequency factor portfolio construction remains under-explored. A key obstacle is the high-dimensional, unbalanced state space created by stocks that enter and exit the investable universe. We introduce Conditional Auto-encoded Factor-based Portfolio Optimisation (CAFPO), which compresses stock-level returns into a small set of latent factors conditioned on 94 firm-specific characteristics. The factors feed a DRL agent implemented with both PPO and DDPG to generate continuous long-short weights. On 20 years of U.S. equity data (2000--2020), CAFPO outperforms equal-weight, value-weight, Markowitz, vanilla DRL, and Fama--French-driven DRL, delivering a 24.6\% compound return and a Sharpe ratio of 0.94 out of sample. SHAP analysis further reveals economically intuitive factor attributions. Our results demonstrate that factor-aware representation learning can make DRL practical for institutional, low-turnover portfolio management.

因子投资强化学习组合优化金融AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。