用深度学习从造血细胞中挖掘隐藏基因特征,精准诊断血液病。
Deep Learning Approaches for Blood Disease Diagnosis Across Hematopoietic Lineages
- 通过自编码器将两万多个基因压缩至256维潜在空间,捕捉分化前后的关键信息。
- 多类分类准确率超95%,零样本预测在二分类任务F1超0.7。
- 适合血液病研究者、生物信息学从业者,可拓展至其他细胞类型分析。
我们提出一种基础建模框架,利用深度学习揭示造血谱系中的潜在遗传特征。该方法在多能祖细胞上训练全连接自编码器,将超过20,000个基因特征降维至256维潜在空间,有效捕捉祖细胞及下游分化细胞(如单核细胞、淋巴细胞)的预测信息。通过训练前馈网络、Transformer和图卷积模型验证嵌入质量,用于血液病诊断任务。同时探索使用祖细胞疾病状态分类模型进行零样本预测,以推断下游细胞状态。模型在多分类任务中准确率超过95%,在零样本二分类任务中F1分数超过0.7。未来工作需进一步优化嵌入以提升对淋巴细胞分类的鲁棒性。
原文摘要 · Abstract (English)
We present a foundation modeling framework that leverages deep learning to uncover latent genetic signatures across the hematopoietic hierarchy. Our approach trains a fully connected autoencoder on multipotent progenitor cells, reducing over 20,000 gene features to a 256-dimensional latent space that captures predictive information for both progenitor and downstream differentiated cells such as monocytes and lymphocytes. We validate the quality of these embeddings by training feed-forward, transformer, and graph convolutional architectures for blood disease diagnosis tasks. We also explore zero-shot prediction using a progenitor disease state classification model to classify downstream cell conditions. Our models achieve greater than 95% accuracy for multi-class classification, and in the zero-shot setting, we achieve greater than 0.7 F1-score on the binary classification task. Future work should improve embeddings further to increase robustness on lymphocyte classification specifically.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。