用强化学习训练大模型,让搜索代理更擅长长周期多轮任务。
Reinforcement Learning for Long-Horizon Multi-Turn Search Agents
- 用强化学习让大模型从经验中学习,提升多轮交互能力。
- 140亿参数模型在法律文档搜索中达到85%准确率,优于前沿模型的78%。
- 允许更长对话轮次能显著提升性能,适合复杂任务场景。
大型语言模型(LLM)代理可通过多轮交互和工具调用解决复杂任务,基于提示的方法已表现出色。本文表明,通过强化学习(RL)从经验中学习,可进一步显著提升能力。在法律文档搜索基准上的实验显示,我们的140亿参数模型在准确率上达到85%,优于前沿模型的78%。此外,我们探索了训练和测试时的轮次限制情形,发现若允许代理在更长的多轮周期中运行,其表现会更好。
原文摘要 · Abstract (English)
Large Language Model (LLM) agents can leverage multiple turns and tools to solve complex tasks, with prompt-based approaches achieving strong performance. This work demonstrates that Reinforcement Learning (RL) can push capabilities significantly further by learning from experience. Through experiments on a legal document search benchmark, we show that our RL-trained 14 Billion parameter model outperforms frontier class models (85% vs 78% accuracy). In addition, we explore turn-restricted regimes, during training and at test-time, that show these agents achieve better results if allowed to operate over longer multi-turn horizons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。