用数学框架评估和改进复杂问题研究智能体的结构能力。
From Intent to Evidence: A Categorical Approach for Structural Evaluation of Deep Research Agents
- 基于范畴论构建研究过程的结构化评估框架。
- 16个前沿系统平均准确率仅19.9%,长程合成与交叉验证仍困难。
- 提出可追踪搜索与显式验证工具,提升系统可靠性。
深度研究智能体(DRAs)通过网络搜索、证据核查与多源信息整合来回答复杂问题。本文提出一种范畴论框架,将深度研究视为从用户意图到证据支撑结论的结构化映射,使检索路径、跨源对齐与最终综合过程显式化。基于此,我们构建了一个包含296个双语问题的机制感知基准,聚焦四项核心研究能力:多跳证据链追踪、跨源声明验证、碎片化信息重组、无效假设拒绝。在人工验证下评估16个前沿系统,发现最佳系统平均准确率仅为19.9%。结果表明,强模型虽能重构证据并识别错误前提,但在长时序合成与高交集验证任务上仍表现不佳。该理论还指导了实际系统改进,如引入可追踪搜索与类别工具,显著提升基于API的深度研究系统性能。本工作既提供挑战性评测基准,也给出可靠研究智能体的设计路径。
原文摘要 · Abstract (English)
Deep Research Agents (DRAs) aim to answer complex questions by searching the web, checking evidence, and synthesizing conclusions across heterogeneous sources. We introduce a category-theoretic framework for evaluating and improving such agents. The framework treats deep research as a structured mapping from user intent to evidence-grounded conclusions, making retrieval traces, cross-source alignment, and final synthesis explicit. Guided by this view, we derive a mechanism-aware benchmark of 296 bilingual questions. The benchmark targets four structural skills central to real research: following multi-hop evidence chains, verifying claims across sources, re-ordering fragmented information, and rejecting unsupported assumptions. We evaluate 16 frontier systems with human verification and find that these structural tasks remain highly challenging: the best system reaches only 19.9% average accuracy. The results show that strong agents can sometimes reorganize evidence and detect false premises, but still struggle with long-horizon synthesis and intersection-heavy verification. Beyond evaluation, the same theory also leads to practical system improvements. We instantiate theory-guided interventions such as tracked search, which preserves retrieval traces, and category tools, which add explicit verification and synthesis steps. These interventions yield measurable gains in API-based deep research systems. Our work therefore provides both a challenging benchmark and concrete design guidance for building more reliable research agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。