挑战真实病理场景下的有丝分裂检测,评估模型跨肿瘤类型与复杂区域的泛化能力。
Mitosis Detection in the Wild: Multi-Tumor and Context-Aware Generalization in the MIDOG 2025 Challenge

- 构建多物种、多肿瘤类型的全切片检测数据集,覆盖真实临床场景的多样性。
- 模型在难点区域误检率翻倍,不同肿瘤类型间性能差异显著,暴露现有模型盲区。
- 引入非典型有丝分裂图象分类,验证集成方法提升性能,但测试时增强无效。
自动化有丝分裂检测是计算病理学中的经典任务。以往基准主要关注扫描仪引起的域偏移,但临床实际应用要求模型具备对组织学景观巨大变异的鲁棒性。MIDOG 2025挑战赛旨在评估算法在前所未有的生物与上下文多样性下的表现。我们构建了包含365例样本的测试集,涵盖12种人类、犬类和猫类肿瘤类型,并在多个扫描平台数字化。挑战不限于人工选定热点区域,还要求在随机组织区域(代表整张切片检测)及高难度区域(富含难分样本)进行检测。第二赛道新增非典型有丝分裂图象(AMFs)分类任务。共18支队伍提交检测赛道,F1最高达0.740;21支队伍参与AMF检测,平衡准确率最高达0.908。分析显示,尽管多数模型在传统热点表现稳定,但在难点区域误检率翻三倍;不同肿瘤类型间性能差异明显,暴露出当前顶尖架构对罕见或高度异型恶性肿瘤的识别盲区。此外,集成方法使F1平均提升1.5个百分点,平衡准确率提升1.3个百分点;而测试时增强(TTA)未带来有效改善。MIDOG 2025表明,‘真实世界’中的有丝分裂检测仍面临重大挑战。从仅热点评估转向多情境框架,为临床可靠性提供了更真实的评价标准。
原文摘要 · Abstract (English)
Automated mitosis detection is a well-established task in computational pathology. While previous benchmarks focused on scanner-induced domain shift, clinical "real-world" application requires models to be robust across the vast variance to be expected in the histological landscape. The MItosis DOmain Generalization (MIDOG) 2025 challenge was designed to evaluate algorithmic performance across unprecedented biological and contextual diversity. We curated a test dataset of 365 cases, encompassing 12 distinct human, canine and feline tumor types, digitized across multiple scanning platforms. Moving beyond hand-selected hotspots, the challenge required detection also in random tissue areas (representative of the whole slide detection situation) and challenging areas (areas rich in hard negatives). In the second track, we introduced the classification of atypical mitotic figures (AMFs). There were 18 teams submitting to the detection track, with F1 scores ranging up to 0.740. In the AMF detection track, we had 21 submissions with balanced accuracy values up to 0.908. Our analysis reveals that while most models perform reliably in traditional hotspots, significant performance degradation occurs in challenging ROIs, where false positive rates tripled. Furthermore, performance varied significantly across the 12 tumor types, highlighting "blind spots" in current state-of-the-art architectures when encountering rare or highly pleomorphic malignancies. Moreover, we evaluated the effectiveness of ensembling and found a mean increases of 1.5 and 1.3 percentage points in F1 score and balanced accuracy, respectively. In contrast, TTA showed no relevant improvement. MIDOG 2025 demonstrates that "in the wild" mitosis detection remains a significant hurdle. The transition from hotspot-only evaluation to a multi-contextual framework provides a more realistic proxy for clinical reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。