arXiv:2607.23524cs.AI2026-07

提出可分解的深度搜索委托智能评估框架,精准诊断模型何时该搜、如何搜。

Delegation Intelligence in Deep Search: A Controllable Framework for Disentangled Capability Diagnosis

论文配图:Delegation Intelligence in Deep Search: A Controllable Framework for Disentangled Capability Diagnosis
图 1 · 摘自论文原文
  • 将委托智能拆解为搜索决策与信息整合验证两维度
  • 构建可控合成流水线,实现多维度能力独立评测
  • 推出DelegSearchBench基准,支持对齐能力的精细化分析

深度搜索正成为现代智能体的核心能力,但现有评估仅依赖最终答案准确率,将检索质量、长上下文理解、证据验证与工具使用决策混杂在一起,难以判断模型是否真正掌握信息委托的时机与方式。为此,本文首次形式化定义深度搜索中的委托智能,并将其分解为两个互补维度:搜索决策(识别信息不足并决定何时、如何搜索)与信息合成与验证(从多源证据中聚合信息,判断来源可靠性,在噪声甚至对抗条件下进行信息整合)。为实现可分离、可复现的测评,我们构建基于文档驱动逆向工程的可控合成流程,提供通用的深度搜索评估构建方法而非单一数据集。作为具体实例,我们建立了DelegSearchBench,并设计了分离各能力维度的评估协议,通过改变文档组成与工具访问条件实现独立测量。在多个代表性模型上的实验表明,仅靠最终答案准确率无法充分刻画深度搜索能力。

原文摘要 · Abstract (English)

Deep search is becoming a core capability of modern agent systems, yet it is typically evaluated solely based on end-to-end answer accuracy. This coupled evaluation paradigm entangles retrieval quality, long-context comprehension, evidence verification, and tool-use decisions, making it difficult to determine whether a model truly knows when and how to delegate information seeking to search. To this end: (1) We formalize this meta-capability as Delegation Intelligence in deep search and decompose it into complementary dimensions-Search Decision-Making (recognizing information insufficiency and deciding whether, when, and how to search) and Information Synthesis and Verification (aggregating evidence from multiple sources, judging source reliability, and synthesizing information under noisy, potentially adversarial conditions). (2) To enable disentangled and reproducible measurement, we develop a controllable synthesis pipeline built on document-grounded reverse engineering. This yields a general recipe for constructing controlled deep-search evaluations rather than a single fixed dataset. (3) As a concrete instantiation, we construct DelegSearchBench, together with a disentangled evaluation protocol that isolates each capability dimension by varying document composition and tool access. (4) Across representative models, we demonstrate that deep-search competence cannot be adequately characterized by final-answer accuracy alone...

深度搜索智能代理评估基准可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。