arXiv:2602.02905cs.AI2026-02被引 3

用真实科研发现测试AI代理,评估其自主探索能力。

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

  • 让AI从问题出发,自主设计实验、编码执行并得出结论。
  • 顶级代理在关键指标上表现有限,成功率低于50%且波动大。
  • 适合研究自主AI科研系统与评估方法的学者使用。

由大语言模型驱动的自主智能体有望实现科学发现的全流程自动化,但对其可验证发现能力的严格评估仍是核心挑战。现有基准存在权衡:要么依赖LLM作为评判者对自动生成的研究成果进行评价,要么优化便捷但孤立的性能指标,这些指标仅粗略反映科学洞察力。为此,我们提出FIRE-Bench(全周期洞察再发现评估),一个通过重新发现近期高影响力机器学习研究中的已验证成果来评估智能体的方法。智能体仅获知来自已发表研究的高层次问题,需自主探索、设计实验、编写代码、执行计划,并基于实证证据得出结论。我们在FIRE-Bench上评估了多种前沿智能体,使用gpt-5等先进模型作为骨干。结果表明,当前智能体在全流程科学研究中仍面临巨大挑战:即使最强代理的再发现成功率也低于50 F1,运行间方差显著,且在实验设计、执行和基于证据的推理中反复出现失败模式。FIRE-Bench为衡量向可靠代理驱动科学发现迈进的进展提供了严谨且具有诊断性的框架。

原文摘要 · Abstract (English)

Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery remains a central challenge. Existing benchmarks face a trade-off: they either heavily rely on LLM-as-judge evaluations of automatically generated research outputs or optimize convenient yet isolated performance metrics that provide coarse proxies for scientific insight. To address this gap, we introduce FIRE-Bench (Full-cycle Insight Rediscovery Evaluation), a benchmark that evaluates agents through the rediscovery of established findings from recent, high-impact machine learning research. Agents are given only a high-level research question extracted from a published, verified study and must autonomously explore ideas, design experiments, implement code, execute their plans, and derive conclusions supported by empirical evidence. We evaluate a range of state-of-the-art agents with frontier LLMs backbones like gpt-5 on FIRE-Bench. Our results show that full-cycle scientific research remains challenging for current agent systems: even the strongest agents achieve limited rediscovery success (<50 F1), exhibit high variance across runs, and display recurring failure modes in experimental design, execution, and evidence-based reasoning. FIRE-Bench provides a rigorous and diagnostic framework for measuring progress toward reliable agent-driven scientific discovery.

AI科研智能体评估科学发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。