首份综述梳理58个大模型,助力单细胞生物智能分析
LLM4Cell: A Survey of Large Language and Agentic Models for Single-Cell Biology
- 分类58个基础与智能体模型,覆盖多组学与空间数据
- 基于40+数据集评估,揭示模型在解释性与公平性短板
- 适合生物信息、AI交叉研究者参考,推动可信赖模型发展
大型语言模型(LLMs)与新兴智能体框架正逐步改变单细胞生物学,实现自然语言推理、生成式注释和多模态数据整合。然而进展仍分散于不同数据模态、架构与评估标准之间。本文首次系统综述了58个专为单细胞研究开发的基础模型与智能体模型,涵盖RNA、ATAC、多组学及空间模态。将这些方法分为五大类:基础型、文本桥梁型、空间型、多模态型、表观基因组型与智能体型,并映射至八项关键分析任务,包括注释、轨迹建模、扰动预测与药物响应预测。基于超过40个公开数据集,分析了基准适用性、数据多样性及伦理与可扩展性约束,并在10个领域维度上评估模型表现,涵盖生物学合理性、多组学对齐、公平性、隐私保护与可解释性。通过关联数据集、模型与评估维度,本工作首次提供语言驱动的单细胞智能全景视图,指出可解释性、标准化与可信模型开发中的开放挑战。
原文摘要 · Abstract (English)
Large language models (LLMs) and emerging agentic frameworks are beginning to transform single-cell biology by enabling natural-language reasoning, generative annotation, and multimodal data integration. However, progress remains fragmented across data modalities, architectures, and evaluation standards. LLM4Cell presents the first unified survey of 58 foundation and agentic models developed for single-cell research, spanning RNA, ATAC, multi-omic, and spatial modalities. We categorize these methods into five families-foundation, text-bridge, spatial, multimodal, epigenomic, and agentic-and map them to eight key analytical tasks including annotation, trajectory and perturbation modeling, and drug-response prediction. Drawing on over 40 public datasets, we analyze benchmark suitability, data diversity, and ethical or scalability constraints, and evaluate models across 10 domain dimensions covering biological grounding, multi-omics alignment, fairness, privacy, and explainability. By linking datasets, models, and evaluation domains, LLM4Cell provides the first integrated view of language-driven single-cell intelligence and outlines open challenges in interpretability, standardization, and trustworthy model development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。