用新骨干网络提升乳腺影像多任务检测模型性能
Modern Backbones Improve Multi-task DETR for Mammography Classification and Lesion Localization

- 采用多任务DETR框架,共享特征同时完成良恶性判断与病灶定位
- ConvNeXtV2在OPTIMAM上达97.96% AUC,DINOv3在SGM1k上表现最优
- 新骨干网络显著优于旧版ResNet,适合临床辅助诊断研究者参考
联合图像级恶性预测与候选区域定位可增强人工智能在乳腺摄影中的实用性。本文基于多任务DETR框架,利用共享表示同时实现图像级恶性判断与病灶定位,在OPTIMAM和经活检确认的SGM1k数据集上进行评估。两个数据集上,现代骨干网络均优于传统ResNet型特征,其中ConvNeXtV2和DINOv3表现最佳,而MambaVision相对较弱。在OPTIMAM上,ConvNeXtV2取得最高综合性能:97.96% AUC,99.89% 敏感性,25.08% [email protected],74.38% [email protected];在SGM1k上,DINOv3表现最强,达到90.97% AUC,86.28% 敏感性,82.00% 特异性,27.04% [email protected],77.32% [email protected]。结果表明,骨干网络质量是多任务乳腺影像分析的关键,ConvNeXtV2在该框架中表现出色,是适配乳腺影像的优秀卷积网络选择。
原文摘要 · Abstract (English)
Joint exam-level prediction and candidate-region localization may improve the usefulness of AI support in mammography. We study this setting using a multi-task DETR framework, where shared representations support both image-level malignancy prediction and lesion localization, and evaluate its performance on OPTIMAM and a biopsy-confirmed SGM1k cohort. Across both datasets, modern backbones consistently outperformed older ResNet-style features, with ConvNeXtV2 and DINOv3 giving the strongest overall results, whereas MambaVision was less competitive. On OPTIMAM, ConvNeXtV2 achieved the best overall performance, reaching 97.96% AUC, 99.89% sensitivity, 25.08% [email protected], and 74.38% [email protected]. On SGM1k, DINOv3 gave the strongest overall results, with 90.97% AUC, 86.28% sensitivity, 82.00% specificity, 27.04% [email protected], and 77.32% [email protected]. These findings suggest that backbone quality is a critical factor in effective multi-task mammography, with ConvNeXtV2 emerging as a particularly strong and well-matched CNN backbone for mammography in this framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。