用主动学习+自监督提升稀有巨核细胞分类准确率
ActiveSSF: An Active-Learning-Guided Self-Supervised Framework for Long-Tailed Megakaryocyte Classification
- 结合高斯滤波与临床先验,精准定位细胞区域
- 动态调整相似度阈值,显著提升罕见亚型识别率
- 适合医学图像分析与稀有细胞分类研究者
精确分类巨核细胞对诊断骨髓增生异常综合征至关重要。尽管自监督学习在医学图像分析中展现出潜力,但其在染色切片中巨核细胞分类仍面临三大挑战:(1) 普遍存在的背景噪声掩盖细胞细节;(2) 长尾分布导致稀有亚型数据不足;(3) 复杂形态变异引发高类内差异。为此,我们提出ActiveSSF框架,融合主动学习与自监督预训练。具体方法包括:利用高斯滤波结合K均值聚类与HSV分析(辅以临床先验知识)实现精准感兴趣区域提取;设计自适应样本选择机制,动态调整相似度阈值以缓解类别不平衡;基于标注样本的原型聚类以应对形态复杂性。在临床巨核细胞数据集上的实验表明,ActiveSSF不仅达到当前最优性能,且显著提升了稀有亚型的识别准确率。该技术集成进一步凸显了其在临床场景中的实际应用潜力。
原文摘要 · Abstract (English)
Precise classification of megakaryocytes is crucial for diagnosing myelodysplastic syndromes. Although self-supervised learning has shown promise in medical image analysis, its application to classifying megakaryocytes in stained slides faces three main challenges: (1) pervasive background noise that obscures cellular details, (2) a long-tailed distribution that limits data for rare subtypes, and (3) complex morphological variations leading to high intra-class variability. To address these issues, we propose the ActiveSSF framework, which integrates active learning with self-supervised pretraining. Specifically, our approach employs Gaussian filtering combined with K-means clustering and HSV analysis (augmented by clinical prior knowledge) for accurate region-of-interest extraction; an adaptive sample selection mechanism that dynamically adjusts similarity thresholds to mitigate class imbalance; and prototype clustering on labeled samples to overcome morphological complexity. Experimental results on clinical megakaryocyte datasets demonstrate that ActiveSSF not only achieves state-of-the-art performance but also significantly improves recognition accuracy for rare subtypes. Moreover, the integration of these advanced techniques further underscores the practical potential of ActiveSSF in clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。