arXiv:2607.20926cs.AI2026-07ACL被引 1

评测大模型在真实科研流程中的信息搜寻与推理能力

SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration

论文配图:SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration
图 1 · 摘自论文原文
  • 构建多任务科研工作流评估基准,覆盖数据库导航到知识融合
  • 复杂任务下模型准确率大幅下降,合成类任务错误率超90%
  • 适合关注AI辅助科研、智能代理评估的研究者使用

科学研究涉及跨异构数据源的复杂信息获取与推理流程。现有基准主要聚焦通用检索或静态科学问答,难以评估真实科研中所需的关键能力。我们提出SciExplore,一个用于评估大语言模型与自主代理在科学信息获取与推理方面能力的基准。该基准包含四类任务,覆盖103个专家精心设计的任务,涵盖十余个科学领域:科学数据库导航、模糊文献检索、缺失参考补全、跨源结构化知识融合,逐步考察从实体级推理、文档级识别到证据级定位和领域级综合等更高阶能力。我们在超过十种前沿大模型与自主代理上进行了评估,发现性能随任务复杂度显著下降,尤其在最复杂的结构化合成任务中准确率极低。结果凸显当前模型与代理在真实科研信息获取场景下的严重局限。

原文摘要 · Abstract (English)

Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily emphasize general-domain retrieval or static scientific question answering, and therefore fail to assess key capabilities required in realistic scientific research workflows. We introduce SciExplore, a benchmark designed to evaluate scientific information-seeking and reasoning capabilities of LLMs and agents. SciExplore comprises four task types covering 103 expert-curated tasks across more than ten scientific disciplines: scientific database navigation, ambiguous literature retrieval, missing reference completion, and cross-source structured knowledge synthesis, which probe progressively higher-level abilities from entity-level reasoning and document-level identification to evidence-level grounding and domain-level synthesis. We evaluate over ten state-of-the-art LLMs and autonomous agents on SciExplore, revealing substantial performance gaps with performance degrading sharply as task complexity increases and extremely low accuracy on the most challenging structured synthesis tasks. These results highlight significant limitations of current models and agents in realistic scientific information-seeking scenarios.

科研智能信息检索智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。