测试大模型对细胞生物学的理解能力,发现现有模型表现远未达预期。
CellVerse: Do Large Language Models Really Understand Cell Biology?
- 构建统一问答基准CellVerse,涵盖细胞、药物、基因三层次分析任务。
- 14个大模型在药物响应预测任务中均不如随机猜测,表现差强人意。
- 揭示大模型在单细胞生物理解上仍存重大挑战,适合生物信息学研究者参考。
近期研究证明可将单细胞数据建模为自然语言,并利用大语言模型(LLM)理解细胞生物学。然而,对LLM在语言驱动的单细胞分析任务中的表现仍缺乏全面评估。为此,我们提出CellVerse,一个以语言为中心的统一问答基准,整合四种单细胞多组学数据,涵盖三个层级的任务:细胞类型注释(细胞级)、药物响应预测(药物级)和扰动分析(基因级)。我们系统评估了14个开源与闭源的LLM(参数量从160M到671B),结果表明:现有专用模型(C2S-Pythia)在所有子任务中均无法做出合理判断;而通用模型如Qwen、Llama、GPT和DeepSeek系列展现出初步理解能力。但总体性能远低于预期,尤其在广泛研究的药物响应预测任务中,所有评估模型均未显著优于随机猜测。CellVerse首次提供了大规模实证证据,表明将LLM应用于细胞生物学仍面临巨大挑战。本工作为通过自然语言推动细胞生物学发展奠定基础,有望促进下一代单细胞分析范式。
原文摘要 · Abstract (English)
Recent studies have demonstrated the feasibility of modeling single-cell data as natural languages and the potential of leveraging powerful large language models (LLMs) for understanding cell biology. However, a comprehensive evaluation of LLMs' performance on language-driven single-cell analysis tasks still remains unexplored. Motivated by this challenge, we introduce CellVerse, a unified language-centric question-answering benchmark that integrates four types of single-cell multi-omics data and encompasses three hierarchical levels of single-cell analysis tasks: cell type annotation (cell-level), drug response prediction (drug-level), and perturbation analysis (gene-level). Going beyond this, we systematically evaluate the performance across 14 open-source and closed-source LLMs ranging from 160M to 671B on CellVerse. Remarkably, the experimental results reveal: (1) Existing specialist models (C2S-Pythia) fail to make reasonable decisions across all sub-tasks within CellVerse, while generalist models such as Qwen, Llama, GPT, and DeepSeek family models exhibit preliminary understanding capabilities within the realm of cell biology. (2) The performance of current LLMs falls short of expectations and has substantial room for improvement. Notably, in the widely studied drug response prediction task, none of the evaluated LLMs demonstrate significant performance improvement over random guessing. CellVerse offers the first large-scale empirical demonstration that significant challenges still remain in applying LLMs to cell biology. By introducing CellVerse, we lay the foundation for advancing cell biology through natural languages and hope this paradigm could facilitate next-generation single-cell analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。