为科学智能体设计真实研究场景的评估框架,提升评测可信度。
HeurekaBench: A Benchmarking Framework for AI Co-scientist
- 用多LLM半自动流程生成基于真实科研代码的开放研究问题。
- 在单细胞生物学中构建sc-HeurekaBench,验证主流智能体性能。
- 发现批判模块可使开源智能体错误率降低22%,缩小与闭源模型差距。
基于大语言模型的推理模型推动了作为合作者的智能体发展,协助完成多步骤科学分析。然而,评估这些系统极具挑战性,需包含数据处理、解读和生成新见解的真实端到端研究场景。为此,我们提出HeurekaBench框架,通过半自动化流程生成基于真实科学研究及其代码库的探索性、开放式研究问题。每个问题均源自具体论文与代码仓库,并由多个LLM协作提取洞见、生成候选工作流,再与报告结果比对验证。我们在单细胞生物学领域构建sc-HeurekaBench基准,用于对比当前最先进的单细胞智能体。进一步展示该基准在量化分析智能体设计选择上的价值:引入批判模块可使开源LLM智能体在处理不完整响应时错误率降低最高达22%,缩小与闭源模型的差距。总体而言,HeurekaBench为科学智能体提供了严谨、端到端的评估路径,其构建基于真实的科研工作流。
原文摘要 · Abstract (English)
LLM-based reasoning models have enabled the development of agentic systems that act as co-scientists, assisting in multi-step scientific analysis. However, evaluating these systems is challenging, as it requires realistic, end-to-end research scenarios that integrate data analysis, interpretation, and the generation of new insights from the experimental data. To address this limitation, we introduce HeurekaBench, a framework to create benchmarks with exploratory, open-ended research questions for experimental datasets. Each such question is grounded in a scientific study and its corresponding code repository, and is created using a semi-automated pipeline that leverages multiple LLMs to extract insights and generate candidate workflows, which are then verified against reported findings. We instantiate the framework in single-cell biology to obtain sc-HeurekaBench benchmark and use it to compare state-of-the-art single-cell agents. We further showcase the benefits of our benchmark for quantitatively analyzing current design choices in agentic systems. We find that the addition of a critic module can improve ill-formed responses for open-source LLM-based agents by up to 22% and close the gap with their closed-source counterparts. Overall, HeurekaBench sets a path toward rigorous, end-to-end evaluation of scientific agents, grounding benchmark construction in real scientific workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。