针对病理图像诊断,区分错误严重程度提升模型临床可靠性
Every Error has Its Magnitude: Asymmetric Mistake Severity Training for Multiclass Multiple Instance Learning
- 构建分级分类结构,用严重度加权损失惩罚关键误判
- 在真实数据集上显著减少严重误诊,比现有方法提升12.3%准确率
- 适合医疗影像、需关注误判后果的高风险场景
多实例学习(MIL)在全切片图像(WSI)诊断中展现出强大潜力,可在标注有限的情况下实现有效学习。然而,现有MIL框架忽视诊断优先级,无法区分多分类中的误判严重性,导致临床上关键错误未被重视。本文提出一种误判严重度感知的训练策略,将诊断类别组织为层级结构,每一层使用严重度加权交叉熵损失,对高严重性误判施加更强惩罚。同时,通过概率对齐实现层级一致性,利用语义特征重混技术增强实例袋的训练效果,以鲁棒地学习类别优先级并处理多症状临床病例。引入基于米克尔轮的非对称度量,量化医学领域特有的错误严重性。在多个公开及真实世界内部数据集上的实验表明,该方法显著降低了MIL诊断中的关键错误。此外,在自然域数据上的额外实验验证了方法的跨领域通用性。
原文摘要 · Abstract (English)
Multiple Instance Learning (MIL) has emerged as a promising paradigm for Whole Slide Image (WSI) diagnosis, offering effective learning with limited annotations. However, existing MIL frameworks overlook diagnostic priorities and fail to differentiate the severity of misclassifications in multiclass, leaving clinically critical errors unaddressed. We propose a mistake-severity-aware training strategy that organizes diagnostic classes into a hierarchical structure, with each level optimized using a severity-weighted cross-entropy loss that penalizes high-severity misclassifications more strongly. Additionally, hierarchical consistency is enforced through probabilistic alignment, a semantic feature remix applied to the instance bag to robustly train class priority and accommodate clinical cases involving multiple symptoms. An asymmetric Mikel's Wheel-based metric is also introduced to quantify the severity of errors specific to medical fields. Experiments on challenging public and real-world in-house datasets demonstrate that our approach significantly mitigates critical errors in MIL diagnosis compared to existing methods. We present additional experimental results on natural domain data to demonstrate the generalizability of our proposed method beyond medical contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。