通过模拟市场压力下的未来收益,提升投资组合再平衡的稳定性与收益。
Portfolio Reinforcement Learning with Scenario-Context Rollout
- 基于宏观经济条件生成未来多变量收益情景,模拟极端市场变化。
- 在31个美股和ETF组合上,夏普比率最高提升76%,最大回撤降低53%。
- 解决强化学习中的奖励-转移不一致问题,适合量化交易研究者。
市场状态切换会引起分布偏移,导致投资组合再平衡策略性能下降。本文提出宏观条件驱动的情景上下文滚动(SCR),在压力事件下生成合理的次日多变量收益情景。然而,历史数据无法提供若策略不同会如何的结果,导致基于情景的奖励引入了时序差分学习中的奖励-转移不一致,破坏强化学习批评者训练的稳定性。我们分析该不一致并发现其导致混合评估目标。基于此,我们构建由滚动推断出的反事实下一状态,并增强批评者代理的自举目标,从而稳定学习并实现可行的偏差-方差权衡。在31个不同的美国股权和ETF组合的样本外测试中,该方法相比经典及基于强化学习的基线,夏普比率最高提升76%,最大回撤降低最多53%。
原文摘要 · Abstract (English)
Market regime shifts induce distribution shifts that can degrade the performance of portfolio rebalancing policies. We propose macro-conditioned scenario-context rollout (SCR) that generates plausible next-day multivariate return scenarios under stress events. However, doing so faces new challenges, as history will never tell what would have happened differently. As a result, incorporating scenario-based rewards from rollouts introduces a reward--transition mismatch in temporal-difference learning, destabilizing RL critic training. We analyze this inconsistency and show it leads to a mixed evaluation target. Guided by this analysis, we construct a counterfactual next state using the rollout-implied continuations and augment the critic agent's bootstrap target. Doing so stabilizes the learning and provides a viable bias-variance tradeoff. In out-of-sample evaluations across 31 distinct universes of U.S. equity and ETF portfolios, our method improves Sharpe ratio by up to 76% and reduces maximum drawdown by up to 53% compared with classic and RL-based portfolio rebalancing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。