评测AI在单细胞数据中推导复杂生物学结论的能力
scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology
- 构建长时序单细胞生物分析基准,要求从原始数据自主推导结论
- 覆盖5类重大生物问题,16/63次运行中最强模型达标(25.4%)
- 适合开发可验证的生物智能系统的研究者使用
单细胞研究需分析师通过多步骤流程和元数据整合,将原始测量转化为具体生物学结论。现有AI-生物基准多聚焦于广义知识或局部分析步骤。我们提出scBench-Long,一个面向长时序单细胞生物学的可验证基准,要求智能体在无预设方法的情况下,从原始或近原始数据中恢复科学结论。该基准包含21项评估,涵盖黑色素瘤CD8 T细胞反应性、CD8 RNA+ATAC调控推断、人猴嵌合体发育、KRAS驱动肺肿瘤衰老及致命新冠肺部病理等。任务涉及配对scRNA/TCR测序、转录组与染色质分析、跨物种转录组学、组合scRNA-seq、单核RNA-seq、免疫受体谱、直系同源图谱、配体-受体资源及验证证据。候选结论经重现、评审并转换为可控答案词汇表,支持确定性评分与轨迹评价标准。在1,068条完成轨迹中,表现最佳的模型-提示组合仅在63次运行中有16次通过(25.4%)。scBench-Long评估智能体是否能超越局部分析,基于单细胞数据做出复杂且有依据的科学推断。
原文摘要 · Abstract (English)
Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay context, and auxiliary evidence. Existing AI-biology benchmarks largely measure broad knowledge, executable workflows, or local analysis steps. We introduce scBench-Long, a benchmark for long-horizon single-cell biology in which agents must recover scientific conclusions from raw or near-raw data without prescribed methods. The benchmark contains 21 evaluations spanning melanoma CD8 T-cell reactivity, CD8 RNA+ATAC regulatory inference, human--monkey chimera development, KRAS-driven lung tumor aging, and lethal COVID-19 lung pathology. Tasks cover paired scRNA/TCR sequencing, RNA and chromatin profiling, cross-species transcriptomics, combinatorial scRNA-seq, single-nucleus RNA-seq, immune repertoires, ortholog maps, ligand--receptor resources, and validation evidence. Candidate claims are reproduced, reviewed, and converted into controlled answer vocabularies with deterministic grading and trajectory rubrics. Across 1,068 completed trajectories, the strongest model--harness pair passes 16/63 runs (25.4\%). scBench-Long evaluates whether agents can move beyond local analysis steps and make complex scientific claims that are supported by single-cell data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。