让AI搜索更智能:通过结构化工具提升深度研究准确性
EigentSearch-Q+: Enhancing Deep Research Agents with Structured Reasoning Tools

- 引入Q+工具链,规划查询、监控进度、提取长网页证据
- 在4个基准上提升准确率,最高达3.8个百分点
- 适合需要可靠网络推理的AI研究与开发人员
深度研究需要对网络证据进行推理以回答开放问题,是智能体的核心能力。然而,许多研究智能体仍依赖隐式、非结构化的搜索行为,导致重复探索和脆弱的证据整合。受Anthropic“think”工具范式及信息检索研究启发,我们提出Q+,一套查询与证据处理工具,使网络搜索更具目的性,包括引导查询规划、监控搜索进展,并从长网页快照中提取证据。我们将Q+集成到Eigent(一个开源、可生产级的多智能体工作流系统)的浏览器子智能体中,形成EigentSearch-Q+。在四个基准测试(SimpleQA-Verified、FRAMES、WebWalkerQA、XBench DeepSearch)中,Q+分别使GPT-4.1、GPT-5.1、Minimax M2.5后端的浏览器智能体基准加权平均准确率提升3.0、3.8和0.6个百分点。案例分析表明,EigentSearch-Q+通过显式化搜索进程与证据处理,生成更连贯的工具调用轨迹。
原文摘要 · Abstract (English)
Deep research requires reasoning over web evidence to answer open-ended questions, and it is a core capability for AI agents. Yet many deep research agents still rely on implicit, unstructured search behavior that causes redundant exploration and brittle evidence aggregation. Motivated by Anthropic's "think" tool paradigm and insights from the information-retrieval literature, we introduce Q+, a set of query and evidence processing tools that make web search more deliberate by guiding query planning, monitoring search progress, and extracting evidence from long web snapshots. We integrate Q+ into the browser sub-agent of Eigent, an open-source, production-ready multi-agent workforce for computer use, yielding EigentSearch-Q+. Across four benchmarks (SimpleQA-Verified, FRAMES, WebWalkerQA, and XBench DeepSearch), Q+ improves Eigent's browser agent benchmark-size-weighted average accuracy by 3.0, 3.8, and 0.6 percentage points (pp) for GPT-4.1, GPT-5.1, and Minimax M2.5 model backends, respectively. Case studies further suggest that EigentSearch-Q+ produces more coherent tool-calling trajectories by making search progress and evidence handling explicit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。