arXiv:2607.10960cs.LGq-fin.CP2026-07

用闭环模拟器研究动态费率下交易执行,发现强化学习可显著降低滑点。

Reinforcement Learning for Execution under Dynamic Fees in a Closed-Loop DEX Simulator

论文配图:Reinforcement Learning for Execution under Dynamic Fees in a Closed-Loop DEX Simulator
图 1 · 摘自论文原文
  • 构建动态费率下的闭环交易模拟环境,包含套利机制与噪声订单流。
  • DQN策略相比基准减少13.3个基点的执行损失,且仅在动态费率下有效。
  • 适合关注AMM执行优化与强化学习应用的研究者。

面向交易者的动态费率正被提议用于自动化做市商(AMMs),但历史数据无法揭示订单流如何响应:交易费用不变化、交易者类型不可观测,回放交易记录也不构成序列决策环境。为此,我们构建了一个最小闭环模拟器,通过构造使缺失信号显式存在:两个由均衡启发的动态费率规则重定价的恒定乘积池、对费率敏感的噪声订单流,以及闭式解的中心化交易所-去中心化交易所套利模型。均衡作为闭合原则,而非交易者需学习的目标。在对比经过调优的阶梯式、规划型、前瞻型及表格型策略后,唯一表现有效的策略是小型DQN,其相对于单步路由策略的改进在所有测试条件下均显著优于零。在预留的最后1000个种子中,强制所有策略完成度为1.0,其在每种步骤内排序下均降低实施滑点,尤其在预设的代理最后排序下,实现13.3个基点的订单面值减损。该优势集中于并源自动态费率环境;在固定费率下,配对差异不可区分地接近零。结果提供关于AMM执行控制的模型条件反事实证据,而非关于历史交易者、均衡博弈或可部署利润的证据。

原文摘要 · Abstract (English)

Trader-facing dynamic fees are increasingly proposed for automated market makers (AMMs), but historical data do not identify how order flow would respond: trader-facing fees do not vary, trader types are latent, and a replayed tape is not a sequential decision environment. We therefore construct a minimal closed-loop simulator in which the missing signal exists by construction: two constant-product pools repriced by an equilibrium-inspired dynamic-fee rule, fee-sensitive noise flow, and closed-form CEX--AMM arbitrage. Equilibrium is used as a closure principle, not as an object the trader learns. Against a tuned benchmark ladder of schedule, planning, lookahead, and tabular policies, a small DQN is the only evaluated valid policy whose paired improvement over tuned one-step routing excludes zero. On a reserved final block of 1{,}000 seeds with completion forced to 1.0 for every policy, it reduces implementation shortfall under every tested intra-step ordering, by $13.3\bps$ of order notional under the pre-specified agent-last ordering, and the edge is concentrated in, and learned from, dynamic-fee environments: under constant fees the paired difference is indistinguishable from zero. The result is model-conditioned counterfactual evidence about execution control in AMMs, not evidence about historical traders, equilibrium play, or deployable profit.

强化学习AMM执行优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。