用生物先验知识构建细胞图谱,提升单细胞数据分析的准确性和效率
DOGMA: Weaving Structural Information into Data-centric Single-cell Transcriptomics Analysis
- 基于多层级生物先验构建细胞关系图,融合细胞本体与基因本体
- 零样本跨物种识别准确率高,显存占用降低70%以上
- 适合需要跨物种分析和低资源部署的研究者
近期数据驱动的人工智能方法已成为单细胞转录组分析的主流范式,将数据表征而非模型复杂度视为核心瓶颈。现有方法大多将细胞视为独立实体,直接应用通用机器学习模型处理其原始序列数据,但忽略了由生物系统功能机制和原始测序数据质量问题驱动的潜在细胞间关系。为此,一系列结构化方法应运而生,虽通过启发式规则捕捉复杂细胞关系并增强原始数据,却常忽视生物先验知识,导致计算开销大、图表示效果不佳。为此,我们提出 DOGMA,一个面向数据结构重塑与语义增强的数据中心框架。该框架通过多层次生物先验知识,实现从统计对齐到细胞本体与系统发育结构的整合,构建生物学可解释的细胞图,并利用基因本体弥补特征层面的语义鸿沟。在复杂的多物种、多器官基准测试中,DOGMA 在严格零样本细胞类型评估中表现出强鲁棒性与高样本效率,且下游推理的显存占用和推理时间显著降低。
原文摘要 · Abstract (English)
Recently, data-centric AI methodology has been a dominant paradigm in single-cell transcriptomics analysis, which treats data representation rather than model complexity as the fundamental bottleneck. In the review of current studies, earlier sequence methods treat cells as independent entities and adapt prevalent ML models to analyze their directly inherited sequence data. Despite their simplicity and intuition, these methods overlook the latent intercellular relationships driven by the functional mechanisms of biological systems and the inherent quality issues of the raw sequencing data. Therefore, a series of structured methods has emerged. Although they employ various heuristic rules to capture intricate intercellular relationships and enhance the raw sequencing data, these methods often neglect biological prior knowledge. This omission incurs substantial overhead and yields suboptimal graph representations, hindering the utility of ML models. To address these issues, we propose DOGMA, a data-centric framework designed for the structural reshaping and semantic enhancement of raw data through multi-level biological prior knowledge. Transcending reliance on purely data-driven heuristics, DOGMA provides a prior-guided graph construction pipeline that integrates statistical alignment with Cell Ontology and phylogenetic structure for biologically grounded cell-graph construction and robust cross-species alignment. Furthermore, Gene Ontology is utilized to bridge the feature-level semantic gap by incorporating functional priors. In complex multi-species and multi-organ benchmarks, DOGMA exhibits strong robustness in strict zero-shot cell-type evaluation and sample efficiency while using substantially lower GPU memory and inference time in downstream evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。