arXiv:2602.09063q-bio.GNcs.AI2026-02被引 2

测试AI在真实单细胞数据上的分析能力,发现模型表现受平台影响极大。

scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis

  • 构建394个真实分析任务的基准,涵盖六种测序平台和七类任务
  • 八款前沿模型准确率仅29%-53%,平台差异导致性能下降超40个百分点
  • 适合开发单细胞分析AI的团队用来诊断模型缺陷

随着单细胞RNA测序数据应用日益广泛、规模与复杂性持续增长,数据分析仍成为多数研究组的瓶颈。尽管前沿AI代理在软件工程和通用数据分析方面大幅提升,但其能否从混乱的现实单细胞数据中提取生物学洞见尚不明确。我们提出scBench,一个由394个可验证问题组成的基准,源自覆盖六种测序平台和七类任务的实际scRNA-seq工作流程。每个问题提供分析步骤前的数据快照,并配备确定性评分器,评估关键生物结果的恢复情况。对八款前沿模型的基准测试显示,准确率在29%-53%之间,且存在显著的模型-任务与模型-平台交互效应。平台选择对准确率的影响堪比模型选择,低文档化技术上性能下降超过40个百分点。scBench与SpatialBench共同覆盖两大主流单细胞多组学模态,既可用作测量工具,也可作为诊断视角,助力开发能忠实、可重复分析真实scRNA-seq数据的AI代理。

原文摘要 · Abstract (English)

As single-cell RNA sequencing datasets grow in adoption, scale, and complexity, data analysis remains a bottleneck for many research groups. Although frontier AI agents have improved dramatically at software engineering and general data analysis, it remains unclear whether they can extract biological insight from messy, real-world single-cell datasets. We introduce scBench, a benchmark of 394 verifiable problems derived from practical scRNA-seq workflows spanning six sequencing platforms and seven task categories. Each problem provides a snapshot of experimental data immediately prior to an analysis step and a deterministic grader that evaluates recovery of a key biological result. Benchmark data on eight frontier models shows that accuracy ranges from 29-53%, with strong model-task and model-platform interactions. Platform choice affects accuracy as much as model choice, with 40+ percentage point drops on less-documented technologies. scBench complements SpatialBench to cover the two dominant single-cell modalities, serving both as a measurement tool and a diagnostic lens for developing agents that can analyze real scRNA-seq datasets faithfully and reproducibly.

单细胞分析AI代理基准测试scRNA-seq

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。