用自然语言指令直接分析单细胞数据,让科研人员更高效探索基因表达
A Multi-Modal AI Copilot for Single-Cell Analysis with Instruction Following
- 通过文本指令与单细胞数据配对,构建多模态语言模型
- 支持细胞类型注释、伪细胞生成等任务,性能优于现有模型
- 适合生物科研人员快速上手,降低数据分析门槛
大型语言模型擅长理解复杂自然语言指令,可执行多种任务。在生命科学中,单细胞RNA测序(scRNA-seq)数据是细胞生物学的“语言”,记录单个细胞的精细基因表达模式。但传统工具交互低效且不直观,限制研究进展。为此,我们提出InstructCell——一个基于自然语言指令的多模态AI协作者,能直接解析并处理单细胞数据。我们构建了一个包含多种组织与物种scRNA-seq数据的多模态指令数据集,配套开发了可同时理解文本与数据的多模态细胞语言架构。InstructCell使研究人员仅凭自然语言命令即可完成细胞类型注释、条件性伪细胞生成与药物敏感性预测等关键任务。大量评估表明,其性能持续优于或达到现有单细胞基础模型水平,且适应多种实验条件。更重要的是,该工具为复杂单细胞数据探索提供了直观入口,降低了技术门槛,助力更深入的生物学发现。
原文摘要 · Abstract (English)
Large language models excel at interpreting complex natural language instructions, enabling them to perform a wide range of tasks. In the life sciences, single-cell RNA sequencing (scRNA-seq) data serves as the "language of cellular biology", capturing intricate gene expression patterns at the single-cell level. However, interacting with this "language" through conventional tools is often inefficient and unintuitive, posing challenges for researchers. To address these limitations, we present InstructCell, a multi-modal AI copilot that leverages natural language as a medium for more direct and flexible single-cell analysis. We construct a comprehensive multi-modal instruction dataset that pairs text-based instructions with scRNA-seq profiles from diverse tissues and species. Building on this, we develop a multi-modal cell language architecture capable of simultaneously interpreting and processing both modalities. InstructCell empowers researchers to accomplish critical tasks-such as cell type annotation, conditional pseudo-cell generation, and drug sensitivity prediction-using straightforward natural language commands. Extensive evaluations demonstrate that InstructCell consistently meets or exceeds the performance of existing single-cell foundation models, while adapting to diverse experimental conditions. More importantly, InstructCell provides an accessible and intuitive tool for exploring complex single-cell data, lowering technical barriers and enabling deeper biological insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。