评测大模型在论文检索与分类上的能力,发现现有系统表现远低于专家水平。
Can Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert Taxonomies
- 构建基于72篇顶会综述的TaxoBench基准,测试模型端到端检索与结构化组织能力。
- 最佳模型仅召回20.92%专家引用论文,且无一达到专家平均4.86层的分类深度。
- 揭示检索与层级组织是独立瓶颈,需校准评估指标才能公平比较模型性能。
深度研究代理在自动化综述撰写方面日益重要,但现有基准未能同时检验其是否能获取专家认为关键的论文,并将这些论文组织成基于论文的分类体系。我们提出TaxoBench基准,涵盖72篇高被引LLM综述、3,815篇被引论文及其专家撰写的分类体系。该基准在两种设置下评估系统:深研模式测试从主题出发的端到端检索与组织,底向模式提供专家论文集以隔离组织能力。通过ARI和V-Measure评估叶节点分配,通过US-TED、US-NTED和Sem-Path评估层级结构。在7个深度研究代理和16个LLM配置中,最优代理仅召回20.92%专家引用论文,且70个标准底向运行均未达到专家平均4.86层的分类深度。受控实验表明,达到该深度的模型通过碎片化分类实现,降低与专家参考的对齐度。此外,即使新模型在ARI上提升3.68个百分点,原始Sem-Path仍接近无组织基线;在深度匹配后,人类在10个匹配综述上仍领先13.27个百分点。结果表明,检索与层级组织是独立瓶颈,且层级评估指标必须校准后方可用于模型对比。
原文摘要 · Abstract (English)
Deep Research Agents increasingly automate survey writing, yet existing benchmarks do not jointly test whether they retrieve the papers experts consider essential and organize those papers into paper-grounded taxonomies. We introduce TaxoBench, a benchmark built from 72 highly cited LLM surveys, 3,815 cited papers, and their expert-authored taxonomies. TaxoBench evaluates systems in two settings: Deep Research mode measures end-to-end retrieval and organization from a topic, while Bottom-Up mode provides the expert paper set and isolates organization. We evaluate leaf-level assignments with ARI and V-Measure and hierarchy-level structure with US-TED, US-NTED, and Sem-Path. Across 7 Deep Research Agents and 16 LLM configurations, the best agent retrieves only 20.92% of expert-cited papers, and none of 70 standard Bottom-Up runs reaches the experts' average depth of 4.86. A controlled probe shows that models which match this depth do so by fragmenting the taxonomy, reducing alignment with the expert reference. We further find that raw Sem-Path remains near a no-organization floor even when a newer model generation gains 3.68 pp ARI; after depth matching, humans lead on all 10 matched surveys by 13.27 pp. These results identify retrieval and hierarchical organization as separate bottlenecks and show why hierarchy metrics must be calibrated before they are used to compare models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。