arXiv:2510.21652cs.AIcs.CL2025-10被引 39

打造科学研究领域严谨评估AI智能体的基准工具集

AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite

  • 构建包含2400+问题的全流程科研任务库
  • 首次实现生产级搜索工具支持的可复现评估
  • 提供9类科学优化代理与完整基线,适合研究者验证进展

AI智能体有望通过自动化文献综述、实验复现、数据分析甚至提出新研究方向,彻底改变科学生产力。目前已有多种通用型‘深度研究’系统及专用科学智能体(如AI Scientist和AIGS)。然而现有评估基准存在诸多不足:缺乏可复现的智能体工具以控制对比核心能力;未考虑模型成本与工具访问等混淆变量;缺少标准化接口用于快速原型开发与评估;缺乏真实科研场景下的综合评价指标;且缺少全面基线难以识别真正进步。为此,我们提出更严谨评估的原则与工具,构建了AstaBench——首个涵盖完整科研流程的综合性基准套件,包含2400+跨多个科学领域的任务,许多源自实际部署的Asta智能体用户请求。该套件配备首个具备生产级检索工具的科研环境,支持可控、可复现评估,并有效控制混杂因素。同时提供九类科学优化的Asta智能体与大量基线。对57个智能体、22类方法的广泛评估揭示:尽管在某些单项能力上取得进展,但当前AI仍远未解决科研辅助的核心挑战。

原文摘要 · Abstract (English)

AI agents hold the potential to revolutionize scientific productivity by automating literature reviews, replicating experiments, analyzing data, and even proposing new directions of inquiry; indeed, there are now many such agents, ranging from general-purpose "deep research" systems to specialized science-specific agents, such as AI Scientist and AIGS. Rigorous evaluation of these agents is critical for progress. Yet existing benchmarks fall short on several fronts: they often (1) lack reproducible agent tools necessary for a controlled comparison of core agentic capabilities; (2) do not account for confounding variables such as model cost and tool access; (3) do not provide standardized interfaces for quick agent prototyping and evaluation; (4) fail to provide holistic, product-informed measures of real-world use cases such as science research; and (5) lack comprehensive baseline agents necessary to identify true advances. In response, we define principles and tooling for more rigorously benchmarking agents. Using these, we present AstaBench, a suite that provides a holistic measure of agentic ability to perform scientific research, comprising 2400+ problems spanning the entire scientific discovery process and multiple scientific domains, and including many problems inspired by actual user requests to deployed Asta agents. Our suite comes with the first scientific research environment with production-grade search tools that enable controlled, reproducible evaluation, better accounting for confounders. Alongside, we provide a comprehensive suite of nine science-optimized classes of Asta agents and numerous baselines. Our extensive evaluation of 57 agents across 22 agent classes reveals several interesting findings, most importantly that despite meaningful progress on certain individual aspects, AI remains far from solving the challenge of science research assistance.

AI智能体科研自动化基准测试可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。