用少量标注实现白血病细胞精准检测与属性分析
Leveraging Sparse Annotations for Leukemia Diagnosis on the Large Leukemia Dataset
- 仅需标记小区域,利用稀疏标注训练模型全视野识别
- 构建48患者来源的大规模白血病数据集LLD,含7类形态属性
- 多任务模型同步检测细胞并预测属性,提升诊断可解释性
白血病是全球第10大常见癌症,也是癌症相关死亡的主要原因。真实诊疗需进行白血球(WBC)定位、分类及形态评估。尽管深度学习在医学影像中取得进展,但缺乏大规模、多样化的多任务数据集;现有小规模数据集缺乏领域多样性,限制实际应用。为克服此问题,我们提出一个名为大型白血病数据集(LLD)的超大规模数据集,并开发了新型白血球检测与属性分析方法。第一,通过48名患者的外周血涂片(PBF),结合多种显微镜、摄像头和放大倍数,构建了大规模数据集,每例白血球在100x下标注7种形态属性(如细胞大小、核形)。第二,提出一个多任务模型,不仅能检测白血球,还能预测其属性,提供可解释且临床有意义的解决方案。第三,提出一种基于稀疏标注的白血球检测与属性分析方法,仅需医生标记视野中一小区域,模型即可利用整个视野信息,提升学习效率与诊断准确率。该数据集与代码已公开。
原文摘要 · Abstract (English)
Leukemia is the 10th most frequently diagnosed cancer and one of the leading causes of cancer-related deaths worldwide. Realistic analysis of leukemia requires white blood cell (WBC) localization, classification, and morphological assessment. Despite deep learning advances in medical imaging, leukemia analysis lacks a large, diverse multi-task dataset, while existing small datasets lack domain diversity, limiting real-world applicability. To overcome dataset challenges, we present a large-scale WBC dataset named Large Leukemia Dataset (LLD) and novel methods for detecting WBC with their attributes. Our contribution here is threefold. First, we present a large-scale Leukemia dataset collected through Peripheral Blood Films (PBF) from 48 patients, through multiple microscopes, multi-cameras, and multi-magnification. To enhance diagnosis explainability and medical expert acceptance, each leukemia cell is annotated at 100x with 7 morphological attributes, ranging from Cell Size to Nuclear Shape. Secondly, we propose a multi-task model that not only detects WBCs but also predicts their attributes, providing an interpretable and clinically meaningful solution. Third, we propose a method for WBC detection with attribute analysis using sparse annotations. This approach reduces the annotation burden on hematologists, requiring them to mark only a small area within the field of view. Our method enables the model to leverage the entire field of view rather than just the annotated regions, enhancing learning efficiency and diagnostic accuracy. From diagnosis explainability to overcoming domain-shift challenges, the presented datasets can be used for many challenging aspects of microscopic image analysis. The datasets, code, and demo are available at: https://im.itu.edu.pk/sparse-leukemiaattri/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。