arXiv:2504.00020q-bio.GNcs.AI2025-04被引 2

针对罕见细胞类型标注难题,提出新型基因组语言模型Celler

Celler:A Genomic Language Model for Long-Tailed Single-Cell Annotation

  • 设计高斯膨胀损失函数动态加权,增强对稀有类别的学习能力
  • 引入困难样本挖掘策略,显著提升少数类别预测准确率
  • 构建含4000万细胞的大型数据集Celler-75,覆盖75种疾病

单细胞技术的进步为解析人类特有疾病等复杂生物系统提供了前所未有的机遇,但也带来了海量长尾分布单细胞数据高效标注的新挑战。为此,我们提出Celler——一种专为单细胞数据标注设计的先进生成预训练模型。该模型包含两项创新:首先,提出高斯膨胀(GInf)损失函数,通过动态调整样本权重,显著增强模型对稀有类别学习的能力,同时降低常见类别的过拟合风险;其次,将困难样本挖掘(HDM)策略引入训练过程,专门针对难以学习的少数类别样本,大幅提高模型预测精度。此外,我们构建了大规模单细胞数据集Celler-75,包含4000万细胞,覆盖80种人体组织和75种特定疾病,为全面探索单细胞技术在疾病研究中的潜力提供关键支持。代码已开源于https://github.com/AI4science-ym/HiCeller。

原文摘要 · Abstract (English)

Recent breakthroughs in single-cell technology have ushered in unparalleled opportunities to decode the molecular intricacy of intricate biological systems, especially those linked to diseases unique to humans. However, these progressions have also ushered in novel obstacles-specifically, the efficient annotation of extensive, long-tailed single-cell data pertaining to disease conditions. To effectively surmount this challenge, we introduce Celler, a state-of-the-art generative pre-training model crafted specifically for the annotation of single-cell data. Celler incorporates two groundbreaking elements: First, we introduced the Gaussian Inflation (GInf) Loss function. By dynamically adjusting sample weights, GInf Loss significantly enhances the model's ability to learn from rare categories while reducing the risk of overfitting for common categories. Secondly, we introduce an innovative Hard Data Mining (HDM) strategy into the training process, specifically targeting the challenging-to-learn minority data samples, which significantly improved the model's predictive accuracy. Additionally, to further advance research in this field, we have constructed a large-scale single-cell dataset: Celler-75, which encompasses 40 million cells distributed across 80 human tissues and 75 specific diseases. This dataset provides critical support for comprehensively exploring the potential of single-cell technology in disease research. Our code is available at https://github.com/AI4science-ym/HiCeller.

单细胞长尾标注生成模型疾病研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。