用语言驱动的分层强化学习优化配对交易,无需微调模型。
Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading

- 高阶抽象与低阶执行均由大语言模型控制,通过提示词优化。
- 在真实市场数据上表现优于传统方法和基线LLM模型。
- 适合研究分层决策、延迟反馈场景下的智能交易系统。
许多序列决策问题具有层次结构,高层语义选择约束底层动作,且反馈延迟且模糊。学习此类问题面临信用分配难题:性能下降可能源于错误的抽象、次优执行或两者交互。我们通过配对交易这一领域研究该挑战,该领域自然融合了资产对选择的长期语义推理与部分可观测条件下的短期执行。我们将配对交易建模为分层强化学习问题,提出一种语言驱动的优化框架,其中高低层策略均由大语言模型参数化,并仅通过提示词更新进行优化。该方法利用预训练的大语言模型作为分层策略,使用轨迹级和回合级文本反馈来适应抽象与执行,无需梯度微调。通过显式分离抽象选择与执行,该框架降低跨层级非平稳性,实现延迟反馈下的精准适配。在真实市场数据上的实验表明,该方法持续优于传统及基于大语言模型的基线,验证了语言驱动分层强化学习的有效性。
原文摘要 · Abstract (English)
Many sequential decision-making problems exhibit hierarchical structure, where high-level semantic choices constrain downstream actions and feedback is delayed and ambiguous. Learning in such settings is challenging due to credit assignment: performance degradation may arise from flawed abstractions, suboptimal execution, or their interaction. We study this challenge through pair trading, a domain that naturally combines long-horizon semantic reasoning for asset pair selection with short-horizon execution under partial observability. We formulate pair trading as a hierarchical reinforcement learning problem and propose a language-driven optimization framework in which both high-level and low-level policies are parameterized by large language models (LLMs) and optimized exclusively through prompt updates. Our approach leverages pretrained LLMs as hierarchical policies and uses trajectory- and episode-level textual feedback to adapt abstractions and execution without gradient-based fine-tuning. By explicitly separating abstraction selection from execution, the framework reduces non-stationarity across hierarchical levels and enables targeted adaptation under delayed feedback. Experiments on real-world market data show consistent improvements over traditional and LLM-based baselines, demonstrating the effectiveness of language-driven hierarchical reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。