用单细胞骨髓图像预测白血病基因突变,模型抗噪能力强。
Predicting Genetic Mutations from Single-Cell Bone Marrow Images in Acute Myeloid Leukemia Using Noise-Robust Deep Learning Models
- 先分癌细胞与非癌细胞,再在疑似癌细胞中预测四种突变
- 尽管标签噪声达20%,突变分类仍达85%准确率
- 适合病理医生和医学AI研究者参考
本研究提出一种鲁棒方法,通过单细胞骨髓图像识别粒细胞母细胞,并预测基因突变,应对标签准确性与数据噪声问题。我们训练了一个二分类器区分白血病(母细胞)与非白血病细胞图像,准确率达90%。为评估泛化能力,将该模型应用于一大型无标签数据集,并由两名血液病理学家验证预测结果,发现白血病与非白血病标签的错误率约为20%。基于此噪声水平,我们在被预测为母细胞的图像上训练了一个四分类模型,以识别特定突变。突变标签仅来自单张切片提取的细胞图像集合。尽管存在肿瘤标签噪声,突变分类模型在四个突变类别上仍达到85%准确率,证明对标签不一致具有强鲁棒性。该研究凸显机器学习模型在处理噪声标签时的有效性,可实现准确且临床相关的突变预测,对血液病理诊断具有重要应用前景。
原文摘要 · Abstract (English)
In this study, we propose a robust methodology for identification of myeloid blasts followed by prediction of genetic mutation in single-cell images of blasts, tackling challenges associated with label accuracy and data noise. We trained an initial binary classifier to distinguish between leukemic (blasts) and non-leukemic cells images, achieving 90 percent accuracy. To evaluate the models generalization, we applied this model to a separate large unlabeled dataset and validated the predictions with two haemato-pathologists, finding an approximate error rate of 20 percent in the leukemic and non-leukemic labels. Assuming this level of label noise, we further trained a four-class model on images predicted as blasts to classify specific mutations. The mutation labels were known for only a bag of cell images extracted from a single slide. Despite the tumor label noise, our mutation classification model achieved 85 percent accuracy across four mutation classes, demonstrating resilience to label inconsistencies. This study highlights the capability of machine learning models to work with noisy labels effectively while providing accurate, clinically relevant mutation predictions, which is promising for diagnostic applications in areas such as haemato-pathology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。