arXiv:2512.00986cs.CL2025-12被引 1

构建首个面向学术深度研究的模块化评测基准,解决评估不全面问题。

ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents

  • 基于10大学科200个真实研究案例,构建结构化评测数据集。
  • 提出双模式评估框架,分别测试任务型代理与基础大模型性能。
  • 发现多源检索与跨领域一致性是当前系统主要短板。

学术论文数量激增催生了自动化深度研究(DR)系统的需求,但其有效评估仍是开放难题。现有基准多聚焦检索环节,忽视高阶规划与推理能力;且偏向通用领域,缺乏对学术核心场景的支持。为此,我们提出ADRA-Bank——一个面向学术深度研究代理的模块化评测基准。该基准基于真实学术文献,包含10大学科、200个实例,涵盖研究与综述论文。我们还设计了ADRA-Eval评估范式,利用学术论文的丰富结构,评估规划、检索与推理三大核心能力。该范式包含端到端评估(针对任务代理)和隔离评估(针对基础大模型作为潜在骨干)。实验显示:当前代理虽具特定优势,但在多源检索与跨领域一致性上表现不佳;提升高层规划能力是释放基础大模型推理潜力的关键。通过揭示可操作的失败模式,ADRA-Bank为研发更可靠的自动学术研究助手提供了诊断工具。

原文摘要 · Abstract (English)

A surge in academic publications calls for automated deep research (DR) systems, but accurately evaluating them is still an open problem. First, existing benchmarks often focus narrowly on retrieval while neglecting high-level planning and reasoning. Second, existing benchmarks favor general domains over the academic domains that are the core application for DR agents. To address these gaps, we introduce ADRA-Bank, a modular benchmark for Academic DR Agents. Grounded in academic literature, our benchmark is a human-annotated dataset of 200 instances across 10 academic domains, including both research and review papers. Furthermore, we propose a modular Evaluation Paradigm for Academic DR Agents (ADRA-Eval), which leverages the rich structure of academic papers to assess the core capabilities of planning, retrieval, and reasoning. It employs two complementary modes: an end-to-end evaluation for \task agents and an isolated evaluation for foundational LLMs as potential backbones. Results reveal uneven capabilities: while agents show specialized strengths, they struggle with multi-source retrieval and cross-field consistency. Moreover, improving high-level planning capability is the crucial factor for unlocking the reasoning potential of foundational LLMs as backbones. By exposing these actionable failure modes, ADRA-Bank provides a diagnostic tool to guide the development of more reliable automatic academic research assistants.

学术研究评测基准大模型评估智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。