arXiv:2602.19419cs.LGq-fin.TR2026-02

用强化学习优化自动做市商的流动性管理,降低操作成本并提升收益。

RAmmStein: Regime Adaptation in Mean-reverting Markets with Stein Thresholds -- Optimal Impulse Control in Concentrated AMMs

  • 基于均值回归模型与深度强化学习,自适应调整流动性策略。
  • 实测净回报率达1.60%,减少85%再平衡次数,节省大量手续费。
  • 适合关注DeFi资本效率与自动化运维的研究者和开发者。

去中心化交易所中的集中流动性提供本质上是一个脉冲控制问题。流动性提供者(LPs)在紧密价格区间内集中资金以最大化交易费收入,同时面临再平衡带来的摩擦成本,包括Gas费和滑点。现有方法多采用启发式或阈值策略,未能充分考虑市场动态。本文将流动性管理建模为最优控制问题,并推导出对应的哈密顿-雅可比-贝尔曼拟变分不等式(HJB-QVI)。提出近似解法RAmmStein,一种结合奥恩斯坦-乌伦贝克过程均值回归速度(theta)等特征的深度强化学习方法。实验表明,该智能体能有效划分行动与静止状态空间。进一步扩展为RAmmStein-Width,通过六动作双深度Q网络联合优化再平衡时机与头寸宽度。基于高频率1Hz Coinbase交易数据(超680万笔交易,1000万TVL,1%默认宽度)的真实环境测试显示,RAmmStein实现1.60%的净收益率,为所有非全知策略中最高;而贪婪策略因Gas成本损失高达-8.4%。值得注意的是,其再平衡频率相比贪婪策略降低85%。RAmmStein-Width自主发现极简策略,仅执行9次再平衡、支出40美元Gas,且在高Gas成本下退化更慢。结果表明,具有周期感知的‘懒惰’策略能显著提升资本效率,避免运营成本侵蚀收益。

原文摘要 · Abstract (English)

Concentrated liquidity provision in decentralized exchanges presents a fundamental Impulse Control problem. Liquidity Providers (LPs) face a non-trivial trade-off between maximizing fee accrual through tight price-range concentration and minimizing the friction costs of rebalancing, including gas fees and swap slippage. Existing methods typically employ heuristic or threshold strategies that fail to account for market dynamics. This paper formulates liquidity management as an optimal control problem and derives the corresponding Hamilton-Jacobi-Bellman quasi-variational inequality (HJB-QVI). We present an approximate solution RAmmStein, a Deep Reinforcement Learning method that incorporates the mean-reversion speed (theta) of an Ornstein-Uhlenbeck process among other features as input to the model. We demonstrate that the agent learns to separate the state space into regions of action and inaction. We further extend the framework with RAmmStein-Width, which jointly optimizes rebalancing timing and position width via a 6-action DDQN. We evaluate the framework using high-frequency 1Hz Coinbase trade data comprising over 6.8M trades on a realistic environment (10M TVL, 1% default width). Experimental results show that RAmmStein achieves a net ROI of 1.60%, the highest among all realistic (non-omniscient) strategies, while greedy strategies lose up to -8.4% to gas costs. Notably, the agent reduces rebalancing frequency by 85% compared to greedy rebalancing. RAmmStein-Width discovers extreme parsimony on its own, executing only 9 rebalances and $40 in gas, and degrades more slowly than all active strategies at elevated gas costs. Our results demonstrate that regime-aware laziness can significantly improve capital efficiency by preserving the returns that would otherwise be eroded by the operational costs.

DeFi强化学习自动做市商资本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。