构建细胞类型标准框架,助力单细胞数据整合与标注
The Cell Ontology in the age of single-cell omics
- 基于标准化术语构建跨物种细胞类型本体
- 支持人类细胞图谱等重大计划的数据标注需求
- 探索大语言模型提升本体更新效率的可能
单细胞组学技术实现了对单个细胞的高分辨率分析,揭示了前所未有的细胞多样性。然而,这些数据的规模与异质性要求强有力的整合与注释框架。细胞本体(Cell Ontology, CL)通过提供标准化、跨物种的细胞类型术语,成为实现数据可发现、可访问、可互操作、可重用(FAIR)原则的核心资源,广泛应用于各类平台与工具中。本文总结了CL在多个平台中的实际应用,并介绍了当前对其内容的扩展工作,包括新增转录组定义的细胞类型,紧密协作于人类细胞图谱(Human Cell Atlas)和脑计划细胞图谱网络(Brain Initiative Cell Atlas Network)以满足其需求。文章还探讨了经典与转录组细胞类型定义的统一挑战,以及标记基因整合与大语言模型(LLMs)在提升CL内容质量与工作流效率方面的未来方向。
原文摘要 · Abstract (English)
Single-cell omics technologies have transformed our understanding of cellular diversity by enabling high-resolution profiling of individual cells. However, the unprecedented scale and heterogeneity of these datasets demand robust frameworks for data integration and annotation. The Cell Ontology (CL) has emerged as a pivotal resource for achieving FAIR (Findable, Accessible, Interoperable, and Reusable) data principles by providing standardized, species-agnostic terms for canonical cell types - forming a core component of a wide range of platforms and tools. In this paper, we describe the wide variety of uses of CL in these platforms and tools and detail ongoing work to improve and extend CL content including the addition of transcriptomic types, working closely with major atlasing efforts including the Human Cell Atlas and the Brain Initiative Cell Atlas Network to support their needs. We cover the challenges and future plans for harmonising classical and transcriptomic cell type definitions, integrating markers and using Large Language Models (LLMs) to improve content and efficiency of CL workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。