用专家策略聚合学习股票交易,提升收益稳定性与风险控制。
Learning Stock Trading Policies via Barycenter-Based Adversarial Inverse Reinforcement Learning
- 通过加权瓦舍斯坦均值聚合多源交易策略,生成稳定伪专家数据
- 在四大股市上超越经典规则与深度强化学习方法,收益更稳
- 结合风险约束函数,确保交易行为符合回撤限制,适合量化交易者
由于奖励延迟、噪声大、探索困难及难以显式施加风险约束,利用强化学习设计有效交易策略仍具挑战。本文提出BRaG框架,基于对抗式逆强化学习,从多种异构专家策略中学习交易行为。该框架采用性能加权的瓦舍斯坦均值聚合专家演示,生成捕捉多元交易风格共性的稳定伪专家表示,用于预训练交易策略,缓解强化学习中的不稳定探索问题。随后使用真实市场奖励对预训练策略进行微调。为实现风险感知决策,引入控制屏障函数,约束动作执行并正则化策略学习以满足回撤上限。在美、英、印、台四大主要股指上评估,所提方法在所有市场均优于经典交易规则和近期深度强化学习方法,且具备更优的风险稳定性。
原文摘要 · Abstract (English)
Designing effective trading strategies using reinforcement learning remains challenging due to delayed and noisy rewards, poor exploration, and the difficulty of enforcing explicit risk constraints. In this work, we propose BRaG, a barycenter-based adversarial inverse reinforcement learning framework for stock trading that learns trading behavior from multiple heterogeneous expert strategies. BRaG aggregates expert demonstrations using a performance-weighted Wasserstein barycenter, yielding a stable pseudo-expert representation that captures shared structure across diverse trading styles. This representation is used to pretrain a trading policy via adversarial imitation learning, which alleviates unstable exploration during reinforcement learning. The pretrained policy is subsequently refined using reinforcement learning with true market rewards. To ensure risk-aware decision-making, BRaG incorporates control barrier functions that constrain action execution and regularize policy learning to satisfy drawdown limits. We evaluate the proposed approach on four major global equity markets, including the US, UK, Indian, and Taiwanese indices. Across all the markets, the proposed approach achieves stronger performance than both classical trading rules and recent deep reinforcement learning methods, while exhibiting more stable risk characteristics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。