arXiv:2510.10909cs.AI2025-10被引 17

构建跨论文科研推理评估基准,测试AI如何用工具整合多篇文献答案

PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature

  • 设计可执行平台支持多工具协同,模拟真实科研场景
  • 顶尖AI agent跨论文推理平均准确率仅38.78%,难题集降至18.47%
  • 适合研究智能科研助手、自动化知识整合的团队使用

理解与推理大规模科学文献是基于大语言模型(LLM)智能体的关键挑战。然而,现有研究多局限于单篇论文的无工具任务,主要因缺乏评估跨论文推理与多工具协同的基准。本文提出PaperArena,一个用于评估LLM智能体在需要整合多篇论文信息并借助外部工具回答问题的基准。给定研究问题,智能体需制定推理计划,与多篇论文交互,并调用合适工具生成有依据的答案。为支持标准化评估,我们提供可执行平台,包含模块化工具环境,涵盖多模态解析、上下文检索和程序化计算。实验表明,即使领先的大模型配合成熟代理流程,平均准确率也仅达38.78%,在难题子集上更低至18.47%。我们还分析了推理轨迹,诊断智能体行为,为社区提供改进与评估更强大科研智能体的洞见。

原文摘要 · Abstract (English)

Understanding and reasoning on the large-scale scientific literature is a crucial touchstone for large language model (LLM) based agents. However, existing works are mainly restricted to tool-free tasks within single papers, largely due to the lack of a benchmark that evaluates cross-paper reasoning and multi-tool orchestration in authentic research scenarios. In this work, we propose PaperArena, a benchmark to evaluate LLM-based agents on questions that require integrating information across multiple papers with the assistance of external tools. Given a research question, agents should formulate a reasoning plan, interact with multiple papers, and invoke appropriate tools to produce a well-grounded answer. To support standardized evaluation, we provide a platform for agent execution, offering a modular tool environment including multimodal parsing, context retrieval, and programmatic computation. Experiments reveal that even the leading LLM powering a well-established agentic workflow achieves merely 38.78% average accuracy, while on the hard subset, accuracy drops to only 18.47%. We also analyze reasoning traces and diagnose agent behavior, providing the community with insights to develop and evaluate more capable scientific agents.

科研智能体跨论文推理工具增强评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。