构建乳腺MRI肿瘤分割与治疗反应预测的标准化评测基准
The MAMA-MIA Challenge: Advancing Generalizability and Fairness in Breast MRI Tumor Segmentation and Treatment Response Prediction
- 采用多中心数据建立统一评估框架,仅用治疗前MRI进行联合建模
- 跨洲际测试显示模型性能差异显著,存在准确率与公平性权衡
- 为提升AI在不同人群中的鲁棒性提供公开数据和评估标准
乳腺癌是全球女性中最常见的恶性肿瘤,也是癌症相关死亡的主要原因。动态对比增强磁共振成像在肿瘤特征分析和新辅助化疗监测中具有核心作用。然而,现有乳腺MRI人工智能模型通常基于异质的数据集、研究人群和评估协议开发与评估,导致结果难以直接比较,限制了对模型跨机构鲁棒性的理解。MAMA-MIA挑战赛旨在通过提供标准化基准,联合评估治疗前MRI的原发肿瘤分割与病理完全缓解预测能力。训练队列包含来自美国多个机构的1,506名患者,评估则在三个独立欧洲中心的外部测试集(共574名患者)上进行,以检验跨大陆和跨机构泛化能力。采用统一评分框架,结合预测性能与年龄、绝经状态、乳腺密度等亚组的一致性。共有26支国际团队参与最终评估。结果显示,在统一外部评估下模型表现差异显著,并揭示总体准确率与亚组公平性之间的权衡。该挑战赛提供了标准化数据集、评估协议及公开资源,推动乳腺癌影像中鲁棒且公平的人工智能系统发展。
原文摘要 · Abstract (English)
Breast cancer is the most frequently diagnosed malignancy among women worldwide and a leading cause of cancer-related mortality. Dynamic contrast-enhanced magnetic resonance imaging plays a central role in tumor characterization and treatment monitoring, particularly in patients receiving neoadjuvant chemotherapy. However, existing artificial intelligence models for breast magnetic resonance imaging are typically developed and evaluated using heterogeneous datasets, study populations, and assessment protocols, making direct comparison difficult and limiting understanding of model robustness across institutions and clinically relevant patient subgroups. The MAMA-MIA Challenge was designed to address these challenges by providing a standardized benchmark for the joint evaluation of primary tumor segmentation and prediction of pathologic complete response using pre-treatment magnetic resonance imaging only. The training cohort comprised 1,506 patients from multiple institutions in the United States, while evaluation was conducted on an external test set of 574 patients from three independent European centers to assess cross-continental and cross-institutional generalization. A unified scoring framework combined predictive performance with subgroup consistency across age, menopausal status, and breast density. Twenty-six international teams participated in the final evaluation phase. Results demonstrate substantial performance variability under a common external evaluation framework and reveal trade-offs between overall accuracy and subgroup fairness. The challenge provides standardized datasets, evaluation protocols, and public resources to promote the development of robust and equitable artificial intelligence systems for breast cancer imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。