arXiv:2604.17931cs.AI2026-04被引 9

用轻量虚拟世界让小模型通过强化学习超越大模型的科研能力

LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent

论文配图:LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
图 1 · 摘自论文原文
  • 构建类真实搜索的轻量虚拟环境,替代真实网络搜索训练
  • 40亿参数模型在GAIA和Xbench上分别达71.3%和78.0%准确率
  • 适合想低成本训练高性能科研智能体的研究者

强化学习(RL)已成为训练基于大语言模型的智能体的有效范式。然而,将智能体强化学习扩展至深度研究任务仍受制于两大耦合挑战:手工构造的合成数据难以激发真正的现实世界搜索能力;而训练过程中依赖真实世界搜索则导致不稳定性与高昂成本,制约了智能体强化学习的可扩展性。LiteResearcher是一个使智能体强化学习可扩展的训练框架:通过构建一个模拟真实搜索动态的轻量虚拟世界,实现持续优化的训练方案,使小型搜索智能体能够超越大规模开源及商用模型(如通义千问DeepResearch和Claude-4.5 Sonnet)。具体而言,在GAIA和Xbench等通用基准测试中,LiteResearcher-4B分别取得71.3%和78.0%的准确率,达到开源模型当前最佳水平,证明了可扩展的强化学习训练是深度研究智能体的关键驱动力。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents. However, scaling agentic RL for deep research remains constrained by two coupled challenges: hand-crafted synthetic data fails to elicit genuine real-world search capabilities, and real-world search dependency during RL training introduces instability and prohibitive cost, which limits the scalability of Agentic RL. LiteResearcher is a training framework that makes Agentic RL scalable: by constructing a lite virtual world that mirrors real-world search dynamics, we enable a continuously improving training recipe that empowers a tiny search agent to outperform large-scale open-source and commercial models (e.g., Tongyi DeepResearch and Claude-4.5 Sonnet). Specifically, on common benchmarks such as GAIA and Xbench, our LiteResearcher-4B achieves open-source state-of-the-art results of 71.3% and 78.0% respectively, demonstrating that scalable RL training is a key enabler for Deep Research Agents.

强化学习科研智能体轻量化训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。