arXiv:2511.13762cs.LGcs.AI2025-11AAAI

为基因增量学习建立首个单细胞转录组基准框架

Gene Incremental Learning for Single-Cell Transcriptomics

  • 将图像增量学习思路迁移至基因数据,构建基因增量学习流程
  • 发现基因遗忘问题,通过适配现有方法显著缓解性能下降
  • 首次提供单细胞转录组基因增量学习的完整评测基准

类作为计算机视觉中的基础单元,在增量学习框架中已得到广泛研究。相比之下,虽在诸多领域扮演关键角色且具有持续增长特性,但对‘标记’(tokens)的增量学习研究仍极为有限。这一研究空白主要源于语言中标记的整体性特征,给其增量学习框架设计带来巨大挑战。为此,本文转向一种特定类型的标记——基因,并基于大规模生物数据集单细胞转录组,构建了基因增量学习的全流程框架并建立了相应的评估体系。我们发现基因增量学习同样存在遗忘问题,因此将现有的类别增量学习方法进行适配以缓解基因记忆衰退。通过大量实验,验证了框架设计与评估的有效性,以及方法改进的实用性。最终,我们提供了单细胞转录组中基因增量学习的完整基准。

原文摘要 · Abstract (English)

Classes, as fundamental elements of Computer Vision, have been extensively studied within incremental learning frameworks. In contrast, tokens, which play essential roles in many research fields, exhibit similar characteristics of growth, yet investigations into their incremental learning remain significantly scarce. This research gap primarily stems from the holistic nature of tokens in language, which imposes significant challenges on the design of incremental learning frameworks for them. To overcome this obstacle, in this work, we turn to a type of token, gene, for a large-scale biological dataset--single-cell transcriptomics--to formulate a pipeline for gene incremental learning and establish corresponding evaluations. We found that the forgetting problem also exists in gene incremental learning, thus we adapted existing class incremental learning methods to mitigate the forgetting of genes. Through extensive experiments, we demonstrated the soundness of our framework design and evaluations, as well as the effectiveness of our method adaptations. Finally, we provide a complete benchmark for gene incremental learning in single-cell transcriptomics.

基因学习单细胞增量学习生物信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。