用梯度强化学习让大模型更懂炒股,提升决策能力。
FLAG-Trader: Fusion LLM-Agent with Gradient-based Reinforcement Learning for Financial Trading
- 用微调大模型做策略网络,结合语言理解与强化学习
- 在真实交易场景中提升收益,同时改善其他金融任务表现
- 适合研究智能投研或交易系统的人参考
经过多模态金融数据微调的大语言模型在各类金融任务中展现出出色的推理能力。然而,在交互式金融市场(如交易)这种需要多步、目标导向的复杂场景中,它们的表现仍显不足,亟需更复杂的智能体机制来优化决策。为此,我们提出 extsc{FLAG-Trader},一种统一架构:将语言处理(通过大模型)与基于梯度的强化学习策略优化相结合。其中,部分微调的大模型充当策略网络,既保留预训练知识,又通过参数高效微调适应金融领域。通过交易收益驱动的策略梯度优化,该框架不仅提升了大模型在交易任务中的表现,还增强了其在其他金融任务上的效果。我们提供了充分的实证证据验证这些改进。
原文摘要 · Abstract (English)
Large language models (LLMs) fine-tuned on multimodal financial data have demonstrated impressive reasoning capabilities in various financial tasks. However, they often struggle with multi-step, goal-oriented scenarios in interactive financial markets, such as trading, where complex agentic approaches are required to improve decision-making. To address this, we propose \textsc{FLAG-Trader}, a unified architecture integrating linguistic processing (via LLMs) with gradient-driven reinforcement learning (RL) policy optimization, in which a partially fine-tuned LLM acts as the policy network, leveraging pre-trained knowledge while adapting to the financial domain through parameter-efficient fine-tuning. Through policy gradient optimization driven by trading rewards, our framework not only enhances LLM performance in trading but also improves results on other financial-domain tasks. We present extensive empirical evidence to validate these enhancements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。