arXiv:2601.22113q-fin.TRcs.LG2026-01

用多样性优化生成多种市场条件下的交易策略,提升执行效果。

Diverse Approaches to Optimal Execution Schedule Generation

  • 采用质量-多样性算法生成针对不同流动性与波动性的专精策略
  • 专精策略在特定环境下性能提升8%-10%,CNN模型比VWAP降低3.1bp滑点
  • 适合研究自适应交易系统或量化策略集成的开发者参考

我们首次将MAP-Elites这一质量-多样性算法应用于交易执行。不同于寻找单一最优策略,该方法生成一组按流动性与波动性条件索引的专精策略组合。各专精策略在其行为领域内实现8%-10%的性能提升,而其他单元则表现下降,提示可通过集成改进的专精策略与基准PPO政策来优化整体表现。结果表明,质量-多样性方法在市场状态自适应执行中具有潜力,但每个行为单元需大量计算资源以在所有市场条件下发展出稳健专精策略。为保障实验完整性,我们构建了一个聚焦订单调度而非战术定位的校准版Gymnasium环境,其瞬时影响模型采用指数衰减与平方根量纲缩放,基于400余只美股数据拟合,样本外$R^2>0.02$。在此环境中,两种近端策略优化架构——含MLP与CNN特征提取器——均显著优于行业基线,其中CNN变体在4,900笔未见样本订单(210亿美元名义金额)上实现2.13 bps到达滑点,远低于VWAP的5.23 bps。这些结果验证了模拟环境的真实性,并为质量-多样性方法提供了强有力的单策略基线。

原文摘要 · Abstract (English)

We present the first application of MAP-Elites, a quality-diversity algorithm, to trade execution. Rather than searching for a single optimal policy, MAP-Elites generates a diverse portfolio of regime-specialist strategies indexed by liquidity and volatility conditions. Individual specialists achieve 8-10% performance improvements within their behavioural niches, while other cells show degradation, suggesting opportunities for ensemble approaches that combine improved specialists with the baseline PPO policy. Results indicate that quality-diversity methods offer promise for regime-adaptive execution, though substantial computational resources per behavioural cell may be required for robust specialist development across all market conditions. To ensure experimental integrity, we develop a calibrated Gymnasium environment focused on order scheduling rather than tactical placement decisions. The simulator features a transient impact model with exponential decay and square-root volume scaling, fit to 400+ U.S. equities with $R^2>0.02$ out-of-sample. Within this environment, two Proximal Policy Optimization architectures - both MLP and CNN feature extractors - demonstrate substantial improvements over industry baselines, with the CNN variant achieving 2.13 bps arrival slippage versus 5.23 bps for VWAP on 4,900 out-of-sample orders ($21B notional). These results validate both the simulation realism and provide strong single-policy baselines for quality-diversity methods.

交易执行强化学习多样性优化策略集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。