首个面向真实单细胞数据的AI科学家评估基准,检验其生物发现能力。
Benchmarking AI scientists for omics data driven biological discovery
- 构建真实单细胞数据上的双任务评测框架,包含15个标注数据集和193道生物学问题。
- AI科学家在细胞类型注释上表现接近人类专家,在科学推理上仍有差距。
- 适合生物学家评估AI工具、研究人员改进自研AI系统使用。
大语言模型推动了人工智能科学家的发展,旨在自主分析生物数据并辅助科学发现。然而,现有评估方法或缺乏真实数据支持,或仅关注预设分析结果,无法反映真实数据驱动的生物学研究。为此,我们提出BAISBench(生物人工智能科学家基准),用于评估AI科学家在真实单细胞转录组数据上的表现。该基准包含两项任务:在15个专家标注的数据集上进行细胞类型注释,以及基于41篇已发表研究中的生物学结论生成的193道多选题进行科学发现推理。我们测试了多个代表性AI科学家,并邀请六名研究生级生物信息学人员作为人类基准完成相同任务。结果显示,当前AI科学家虽尚未实现完全自主发现,但在数据驱动研究中已展现出显著潜力。本基准可有效刻画当前AI科学家的能力与局限,有望成为指导未来开发更强大智能体及帮助生物学家筛选实用工具的重要评估框架。代码与数据集见:https://github.com/EperLuo/BAISBench, https://huggingface.co/datasets/EperLuo/BaisBench。
原文摘要 · Abstract (English)
Recent advances in large language models have enabled the emergence of AI scientists that aim to autonomously analyze biological data and assist scientific discovery. Despite rapid progress, it remains unclear to what extent these systems can extract meaningful biological insights from real experimental data. Existing benchmarks either evaluate reasoning in the absence of data or focus on predefined analytical outputs, failing to reflect realistic, data-driven biological research. Here, we introduce BAISBench (Biological AI Scientist Benchmark), a benchmark for evaluating AI scientists on real single-cell transcriptomic datasets. BAISBench comprises two tasks: cell type annotation across 15 expert-labeled datasets, and scientific discovery through 193 multiple-choice questions derived from biological conclusions reported in 41 published single-cell studies. We evaluated several representative AI scientists using BAISBench and, to provide a human performance baseline, invited six graduate-level bioinformaticians to collectively complete the same tasks. The results show that while current AI scientists fall short of fully autonomous biological discovery, they already demonstrate substantial potential in supporting data-driven biological research. These results position BAISBench as a practical benchmark for characterizing the current capabilities and limitations of AI scientists in biological research. We expect BAISBench to serve as a practical evaluation framework for guiding the development of more capable AI scientists and for helping biologists identify AI systems that can effectively support real-world research workflows. The BAISBench can be found at: https://github.com/EperLuo/BAISBench, https://huggingface.co/datasets/EperLuo/BaisBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。