arXiv:2603.28376cs.CLcs.AI2026-03被引 11

通过验证驱动设计,让小型研究智能体在复杂任务中表现媲美大模型。

Marco DeepResearch: Unlocking Efficient Deep Research Agents via Verification-Centric Design

  • 构建三层次验证机制:数据合成、训练轨迹、推理阶段均引入显式验证
  • 600次工具调用内超越多数80亿参数模型,接近300亿参数模型性能
  • 适合追求高效、低成本智能体研发的研究者与工程团队

深度研究智能体可自主开展开放式探究,整合多源信息检索与多步推理以解决现实问题。为在长周期任务中保持能力,训练与推理阶段的可靠验证至关重要。现有范式的主要瓶颈在于问答数据生成、轨迹构建及推理时扩展缺乏显式验证机制,错误在各阶段传播并降低整体性能。为此,我们提出验证驱动的Marco DeepResearch框架,涵盖三个层面:(1) 问答数据合成:在基于图和基于智能体的问答生成中引入验证机制,控制问题难度并确保答案唯一正确;(2) 轨迹构建:设计验证驱动的轨迹生成方法,将显式验证模式注入训练过程;(3) 推理时扩展:利用Marco DeepResearch自身作为验证器,在推理阶段显著提升难题表现。大量实验表明,该智能体在浏览类挑战基准(如BrowseComp和BrowseComp-ZH)上显著优于多数80亿参数级研究智能体,且在最大600次工具调用限制下,甚至超越或接近多个300亿参数模型(如Tongyi DeepResearch-30B)。

原文摘要 · Abstract (English)

Deep research agents autonomously conduct open-ended investigations, integrating complex information retrieval with multi-step reasoning across diverse sources to solve real-world problems. To sustain this capability on long-horizon tasks, reliable verification is critical during both training and inference. A major bottleneck in existing paradigms stems from the lack of explicit verification mechanisms in QA data synthesis, trajectory construction, and test-time scaling. Errors introduced at each stage propagate downstream and degrade the overall agent performance. To address this, we present Marco DeepResearch, a deep research agent optimized with a verification-centric framework design at three levels: \textbf{(1)~QA Data Synthesis:} We introduce verification mechanisms to graph-based and agent-based QA synthesis to control question difficulty while ensuring answers are unique and correct; \textbf{(2)~Trajectory Construction:} We design a verification-driven trajectory synthesis method that injects explicit verification patterns into training trajectories; and \textbf{(3)~Test-time scaling:} We use Marco DeepResearch itself as a verifier at inference time and effectively improve performance on challenging questions. Extensive experimental results demonstrate that our proposed Marco DeepResearch agent significantly outperforms 8B-scale deep research agents on most challenging benchmarks, such as BrowseComp and BrowseComp-ZH. Crucially, under a maximum budget of 600 tool calls, Marco DeepResearch even surpasses or approaches several 30B-scale agents, like Tongyi DeepResearch-30B.

研究智能体验证机制高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。