arXiv:2606.12736cs.AIcs.LG2026-06被引 5

构建科学任务评测平台,评估AI在真实科研中的表现

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales

论文配图:Benchmarking AI Agents for Addressing Scientific Challenges Across Scales
图 1 · 摘自论文原文
  • 设计交互式评测环境,包含200个分步验证的跨领域科学任务
  • 现有AI能高效完成明确的数据分析流程,但难生成新发现或解决开放问题
  • 揭示常见失败模式,为提升自主性与科学推理能力提供方向

AI代理正被广泛开发以加速科学发现,但其在真实研究场景中的实际能力仍不清晰。现有代理评测缺乏科学工作所需的复杂性、异质性和长程推理支持,而科学任务评测又常将研究简化为静态直接问题,且难以实现交互式评估。为此,我们提出SciAgentArena,一个基于多领域真实科研需求的系统性评测基准。该基准包含约200个任务,具备分步验证机制,并提供交互式、代理无关的评估环境,可评估多种AI代理。实验表明,当前代理在结构清晰的数据分析流程中表现良好,但在生成新见解、持续自主探索及应对开放性研究问题方面表现不一。我们进一步分析了代理的常见失败模式,并指出了提升其可靠性、自主性与科学推理能力的关键方向。完整代码、任务和数据集可通过https://sciagentarena.github.io/获取。

原文摘要 · Abstract (English)

AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood. Existing benchmarks for AI agents rarely capture the complexity, heterogeneity, and extended reasoning required by scientific work, whereas benchmarks for scientific tasks often reduce research to static, direct problems and provide limited support for interactive evaluation. Here, we introduce SciAgentArena, a systematic benchmark for evaluating AI agents in real-world scientific research scenarios drawn from emerging needs across multiple domains. SciAgentArena comprises approximately 200 tasks with stepwise verification and an interactive, agent-agnostic environment for assessing diverse AI agents. Using this benchmark, we find that current agents can contribute effectively to well-specified data-analysis workflows, particularly when the task structure and evaluation criteria are clear. However, their performance remains uneven across scientific contexts: agents struggle to generate genuinely novel insights, sustain self-directed exploration, and formulate robust solutions for open-ended research questions. We further characterize common failure modes across agents and identify opportunities for improving their reliability, autonomy, and scientific reasoning. Together, SciAgentArena provides a practical framework for measuring progress in AI agents for science and for guiding the design of future agents capable of addressing complex scientific challenges. Full codes, tasks, and datasets can be accessed via this link: https://sciagentarena.github.io/.

AI代理科学发现评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。