arXiv:2605.05082eess.IVcs.CV2026-05中稿 · the 18th Internati…

AI模型从超声图像预测乳腺密度,外部验证表现良好。

External Validation of Deep Learning Models for BI-RADS Breast Density Prediction from Ultrasound Images

论文配图:External Validation of Deep Learning Models for BI-RADS Breast Density Prediction from Ultrasound Images
图 1 · 摘自论文原文
  • 用三种深度学习模型在独立数据集上验证乳腺密度预测能力。
  • 极端致密乳腺预测准确率最高(AUROC 0.868–0.899),异质致密型最弱(0.699–0.729)。
  • 模型泛化性强,适合临床辅助评估乳腺癌风险,尤其对致密型乳腺有参考价值。

我们对外部独立队列中的三个深度学习模型(DenseNet121、ViT-B/32、ResNet50)进行了验证,用于从乳腺超声检查中预测乳腺密度。外部验证集包含2,000例超声检查,其中500例为癌症病例(初始BI-RADS 1或2,6个月至10年内确诊),1,500例为按设备厂商和检查年份匹配的阴性对照。性能以患者级别AUROC衡量,分为四类密度:A(脂肪型)、B(散在型)、C(异质型)、D(极度致密型)。作为下游评估,将年龄与AI推导的密度结合,使用Tyrer-Cuzick模型预测10年风险,并与基于年龄和影像报告密度的参考模型比较。所有模型在极度致密乳腺中表现最佳(AUROC 0.868–0.899),脂肪型(0.814–0.838)和散在型(0.764–0.799)表现良好,异质型最差(0.699–0.729)。DenseNet121总体表现最优(微平均AUROC 0.885),内部与外部测试结果相近。风险建模方面,年龄+AI密度组合的AUROC(0.541)略低于年龄+影像报告密度(0.570;p=0.23),差异无统计学意义。结果表明,深度学习模型在不同种族构成的外部数据中具有良好的泛化能力,尽管在极度致密型中表现最佳,但异质型仍具挑战,需针对性优化。

原文摘要 · Abstract (English)

We externally validated three deep learning models (DenseNet121, ViT-B/32, and ResNet50) for predicting mammographic breast density from breast ultrasound exams on an independent cohort. The external validation set comprised 2,000 ultrasound exams, including 500 cancer cases defined by an initial negative exam (BI-RADS 1 or 2) followed by a cancer diagnosis within 6 months to 10 years, and 1,500 negative controls matched by manufacturer and study year. Performance was measured using patient-level AUROC across four density categories: A (fatty), B (scattered), C (heterogeneous), and D (extremely dense). As a downstream assessment, we also evaluated 10-year risk prediction by incorporating age and AI-derived density into the Tyrer-Cuzick model and comparing performance against a reference model using age and mammography-reported density. All three models performed best in extremely dense breasts (AUROC 0.868-0.899), with strong performance in fatty (0.814-0.838) and scattered density (0.764-0.799), and lower performance in heterogeneously dense breasts (0.699-0.729). DenseNet121 achieved the highest overall performance (micro-averaged AUROC 0.885), and performance across categories was comparable between internal and external testing. For risk modeling, age combined with AI-derived density yielded a lower AUROC than age combined with mammography-reported density (0.541 vs. 0.570; p = 0.23), with no statistically significant difference. These findings indicate that deep learning models generalize well to external data with different racial composition for breast density assessment. While performance is strongest in extremely dense breasts, heterogeneously dense remains more challenging, highlighting the need for targeted optimization.

乳腺密度深度学习超声诊断风险预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。