用不确定性评估筛选关键病理图,少标注也能达到顶尖分类效果。
Uncertainty Awareness Enables Efficient Labeling for Cancer Subtyping in Digital Pathology
- 通过计算预测置信度向量,动态识别需标注的关键图像。
- 仅用1-10%精选标注数据,即在基准数据集上达到最先进性能。
- 适合标注资源稀缺的病理诊断场景,提升标注效率与模型精度。
机器学习辅助癌症分型是数字病理学的有前景方向。然而,癌症分型模型需要借助专家标注进行精细训练,以确保预测结果具有可量化的确定性(或不确定性)。为此,我们首次将不确定性感知引入自监督对比学习模型:每轮训练后计算一个证据向量,评估模型对预测结果的置信度,并据此生成不确定性评分。该评分用于筛选最需人工标注的关键图像,实现迭代式精准标注。实验表明,仅使用1%-10%的策略性标注,即可在基准数据集上达到当前最优的癌症分型性能。该方法不仅显著减少了对大规模标注数据的依赖,还提升了分类的精确性与效率,特别适用于标注资源有限的临床环境,为数字病理学的未来发展提供了新路径。
原文摘要 · Abstract (English)
Machine-learning-assisted cancer subtyping is a promising avenue in digital pathology. Cancer subtyping models, however, require careful training using expert annotations so that they can be inferred with a degree of known certainty (or uncertainty). To this end, we introduce the concept of uncertainty awareness into a self-supervised contrastive learning model. This is achieved by computing an evidence vector at every epoch, which assesses the model's confidence in its predictions. The derived uncertainty score is then utilized as a metric to selectively label the most crucial images that require further annotation, thus iteratively refining the training process. With just 1-10% of strategically selected annotations, we attain state-of-the-art performance in cancer subtyping on benchmark datasets. Our method not only strategically guides the annotation process to minimize the need for extensive labeled datasets, but also improves the precision and efficiency of classifications. This development is particularly beneficial in settings where the availability of labeled data is limited, offering a promising direction for future research and application in digital pathology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。