arXiv:2510.24126cs.CL2025-10中稿 · to the First Works…被引 2

用强化学习训练大模型,让搜索代理更擅长长周期多轮任务。

Reinforcement Learning for Long-Horizon Multi-Turn Search Agents

  • 用强化学习让大模型从经验中学习,提升多轮交互能力。
  • 140亿参数模型在法律文档搜索中达到85%准确率,优于前沿模型的78%。
  • 允许更长对话轮次能显著提升性能,适合复杂任务场景。

大型语言模型(LLM)代理可通过多轮交互和工具调用解决复杂任务,基于提示的方法已表现出色。本文表明,通过强化学习(RL)从经验中学习,可进一步显著提升能力。在法律文档搜索基准上的实验显示,我们的140亿参数模型在准确率上达到85%,优于前沿模型的78%。此外,我们探索了训练和测试时的轮次限制情形,发现若允许代理在更长的多轮周期中运行,其表现会更好。

原文摘要 · Abstract (English)

Large Language Model (LLM) agents can leverage multiple turns and tools to solve complex tasks, with prompt-based approaches achieving strong performance. This work demonstrates that Reinforcement Learning (RL) can push capabilities significantly further by learning from experience. Through experiments on a legal document search benchmark, we show that our RL-trained 14 Billion parameter model outperforms frontier class models (85% vs 78% accuracy). In addition, we explore turn-restricted regimes, during training and at test-time, that show these agents achieve better results if allowed to operate over longer multi-turn horizons.

强化学习多轮对话大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。