arXiv:2604.00031q-fin.GNcs.LG2026-04

构建可解释的外汇交易强化学习框架,支持真实成本与复杂动作空间。

Decomposable Reward Modeling and Realistic Environment Design for Reinforcement Learning-Based Forex Trading

论文配图:Decomposable Reward Modeling and Realistic Environment Design for Reinforcement Learning-Based Forex Trading
图 1 · 摘自论文原文
  • 分模块设计奖励架构,11个组件可拆解分析,支持逐步诊断。
  • 在EURUSD上训练获0.765最大夏普比率和57.09%累计收益。
  • 扩大动作空间提升收益但增加换手率,适合关注风险控制的研究者。

将强化学习应用于外汇交易仍面临挑战:必须同时满足真实环境、明确奖励函数和表达性强的动作空间,但多数先前研究依赖简化模拟器、单一标量奖励和受限动作表示,限制了可解释性与实际应用价值。本文提出一个模块化强化学习框架,包含三个紧密集成的组件:一个考虑摩擦的执行引擎,实现严格反前瞻语义,观测在t时刻,执行在t+1时刻,账面估值在t+1时刻,同时纳入滑点、佣金、买卖价差、隔夜融资及保证金触发平仓等真实成本;一个由11个组件构成的可分解奖励架构,权重固定并支持每步诊断日志,便于系统性消融与组件级归因;以及一个含10个离散动作的接口,通过合法动作掩码编码显式交易原语,并施加保证金感知的可行性约束。在EURUSD上的实证评估聚焦于学习动态而非泛化能力,发现奖励交互呈强非单调性,额外惩罚未必改善结果;完整奖励配置达到最高训练夏普比率0.765与累计收益57.09%。扩展动作空间虽提升收益,但也增加换手率并降低夏普比,相较保守的3动作基线表现出收益-活跃度权衡,在固定训练预算下,可扩展变体持续降低回撤,组合配置实现最优终点表现。

原文摘要 · Abstract (English)

Applying reinforcement learning (RL) to foreign exchange (Forex) trading remains challenging because realistic environments, well-defined reward functions, and expressive action spaces must be satisfied simultaneously, yet many prior studies rely on simplified simulators, single scalar rewards, and restricted action representations, limiting both interpretability and practical relevance. This paper presents a modular RL framework designed to address these limitations through three tightly integrated components: a friction-aware execution engine that enforces strict anti-lookahead semantics, with observations at time t, execution at time t+1, and mark-to-market at time t+1, while incorporating realistic costs such as spread, commission, slippage, rollover financing, and margin-triggered liquidation; a decomposable 11-component reward architecture with fixed weights and per-step diagnostic logging to enable systematic ablation and component-level attribution; and a 10-action discrete interface with legal-action masking that encodes explicit trading primitives while enforcing margin-aware feasibility constraints. Empirical evaluation on EURUSD focuses on learning dynamics rather than generalization and reveals strongly non-monotonic reward interactions, where additional penalties do not reliably improve outcomes; the full reward configuration achieves the highest training Sharpe (0.765) and cumulative return (57.09 percent). The expanded action space increases return but also turnover and reduces Sharpe relative to a conservative 3-action baseline, indicating a return-activity trade-off under a fixed training budget, while scaling-enabled variants consistently reduce drawdown, with the combined configuration achieving the strongest endpoint performance.

强化学习外汇交易奖励建模真实环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。