让AI研究代理在对抗性环境中自我反思,提升真实科研能力。
MetaResearcher: Scaling Deep Research via Self-Reflective Reinforcement Learning in Adversarial Virtual Environments
- 构建动态对抗环境,逼AI学会辨别信息真伪和解决时间冲突。
- 设计假设生成与矛盾解析任务,推动AI从查资料转向真正研究。
- 多智能体协作+自我反思奖励,显著减少重复搜索行为。
深度研究代理在自主信息搜集与整合方面已展现强大能力,但其训练仍受限于模拟环境的静态性、仅限事实检索的任务设计,以及基于结果的强化学习效率低下。本文提出MetaResearcher框架,通过四个协同维度实现研究代理的规模化训练:首先引入演化虚拟世界,注入时间动态与对抗性虚假信息,迫使代理发展出来源可信度评估与时间冲突解决能力;其次设计发现导向任务(如假设生成、矛盾解析),超越简单事实检索,推动代理向真实研究行为演进;第三,在GRPO框架中提出自省式元奖励机制,联合优化答案正确性、搜索路径效率、反思深度与工具调用多样性,直接解决先前工作中存在的重复动作循环问题;第四,采用异构多智能体群体架构,包含侦察、过滤与合成三类专用模型,通过协同强化学习习得合作研究策略。基于LiteResearcher基础设施,MetaResearcher训练无需额外API成本,同时在GAIA、Xbench-DS基准上实现性能显著提升,并在对抗条件下展现出更强的认知鲁棒性。本文完整呈现了框架设计、训练方法及计划的实验验证。
原文摘要 · Abstract (English)
Deep research agents have demonstrated remarkable capabilities in autonomous information gathering and synthesis, yet their training remains constrained by the static nature of simulated environments, the limits of fact-retrieval-only task designs, and the inefficiency of outcome-based reinforcement learning. In this work, we propose MetaResearcher, a novel framework that scales deep research agent training across four synergistic dimensions. First, we introduce an Evolving Virtual World that injects temporal dynamics and adversarial misinformation into the training environment, forcing agents to develop source credibility assessment and temporal conflict resolution skills. Second, we design Discovery-Oriented Tasks -- including hypothesis generation and contradiction resolution -- that transcend simple fact retrieval and push agents toward genuine research behaviors. Third, we propose a Self-Reflective Meta-Reward mechanism within the GRPO framework that jointly optimizes for answer correctness, search path efficiency, reflection depth, and tool call diversity, directly addressing the repetitive action loop problem observed in prior work. Fourth, we introduce a Heterogeneous Multi-Agent Swarm architecture comprising specialized Scout, Filter, and Synthesizer models that learn collaborative research strategies through coordinated reinforcement learning. Built upon the LiteResearcher infrastructure, MetaResearcher requires zero marginal API cost for training while targeting substantial improvements in both benchmark performance (GAIA, Xbench-DS) and epistemic robustness under adversarial conditions. We present the complete framework design, training methodology, and planned experimental validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。