arXiv:2601.20843cs.AI2026-01

通过动态计划反思与候选融合,提升复杂研究任务的深度报告生成能力。

Deep Researcher with Sequential Plan Reflection and Candidates Crossover (Deep Researcher Reflect Evolve)

  • 采用序列化计划反思机制,实时调整研究路径,保持全局上下文统一。
  • 在DeepResearch Bench上达到46.21分,优于多个主流研究代理。
  • 适合需要高精度、长链条推理的研究场景或学术写作助手开发。

本文提出一种新型Deep Researcher架构,旨在生成复杂博士级课题的详细研究报告,突破并行扩展范式的局限。系统引入两项核心创新:通过反思实现的序列化研究计划优化,以及候选者交叉算法。前者通过持续回顾进展、动态调整计划,维持统一的全局研究上下文,避免知识孤岛;后者部署多个参数不同的LLM候选者,扩大搜索空间,并融合其成果以生成全面结论。最终通过一次性报告生成,确保内容叙事连贯、事实密度高。基于Gemini 2.5 Pro模型,在全球公认的DeepResearch Bench(含100个博士级任务)上评估,本架构取得46.21分,超越Claude Researcher、Nvidia AIQ Research Assistant、Perplexity Research、Kimi Researcher及Grok Deeper Search等主流研究代理,略高于先前静态版本Static DRA,验证了序列化扩展优于并行自洽范式。

原文摘要 · Abstract (English)

This paper introduces a novel Deep Researcher architecture designed to generate detailed research reports on complex PhD level topics by addressing the inherent limitations of the Parallel Scaling paradigm. Our system utilizes two key innovations: Sequential Research Plan Refinement via Reflection and a Candidates Crossover algorithm. The sequential refinement process is demonstrated as an efficient method that allows the agent to maintain a centralized Global Research Context, enabling it to look back at current progress, reason about the research plan, and intelligently make changes at runtime. This dynamic adaptation contrasts with parallel approaches, which often suffer from siloed knowledge. The Candidates Crossover algorithm further enhances search efficiency by deploying multiple LLM candidates with varied parameters to explore a larger search space, with their findings synthesized to curate a comprehensive final research response. The process concludes with One Shot Report Generation, ensuring the final document is informed by a unified narrative and high fact density. Powered by the Gemini 2.5 Pro model, our Deep Researcher was evaluated on the DeepResearch Bench, a globally recognized benchmark of 100 doctoral level research tasks. Our architecture achieved an overall score of 46.21, demonstrating superior performance by surpassing leading deep research agents such as Claude Researcher, Nvidia AIQ Research Assistant, Perplexity Research, Kimi Researcher and Grok Deeper Search present on the DeepResearch Bench actively running leaderboard. This performance marginally exceeds our previous work, Static DRA, and reinforces the finding that sequential scaling consistently outperforms the parallel self consistency paradigm.

深度研究智能代理大模型应用报告生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。