arXiv:2503.13581eess.IVcs.CV2025-03被引 1

评估商用乳腺X线三维成像模型在不同人群中的表现差异。

Subgroup Performance of a Commercial Digital Breast Tomosynthesis Model for Breast Cancer Detection

  • 在16万+份筛查影像上分层分析模型性能
  • 对非侵袭性癌和致密乳腺组织检出率较低
  • 提醒临床部署前需关注模型潜在偏差

尽管研究表明人工智能模型可提升乳腺筛查效果,但针对商业数字乳腺断层成像(DBT)模型的详细亚组评估仍不足。本研究基于埃默里乳腺影像数据集(EMBED)中163,449例筛查影像,对Lunit INSIGHT DBT模型进行细粒度评估。以162,081例阴性影像为负类,1,368例筛查发现的癌症为正类,开展二分类分析,并按人口统计、影像与病理特征分组,识别潜在差异。模型总体AUC达0.91(95%CI: 0.90-0.92),精确率为0.08(95%CI: 0.08-0.08),召回率为0.73(95%CI: 0.71-0.76)。模型在各人群间表现稳健,但对非侵袭性癌(AUC: 0.85, 95%CI: 0.83-0.87)、钙化灶(AUC: 0.80, 95%CI: 0.78-0.82)及致密型乳腺组织(AUC: 0.90, 95%CI: 0.88-0.91)的检测性能显著偏低。结果提示需加强对模型特性的细致评估,警惕其临床应用中的潜在偏倚。

原文摘要 · Abstract (English)

While research has established the potential of AI models for mammography to improve breast cancer screening outcomes, there have not been any detailed subgroup evaluations performed to assess the strengths and weaknesses of commercial models for digital breast tomosynthesis (DBT) imaging. This study presents a granular evaluation of the Lunit INSIGHT DBT model on a large retrospective cohort of 163,449 screening mammography exams from the Emory Breast Imaging Dataset (EMBED). Model performance was evaluated in a binary context with various negative exam types (162,081 exams) compared against screen detected cancers (1,368 exams) as the positive class. The analysis was stratified across demographic, imaging, and pathologic subgroups to identify potential disparities. The model achieved an overall AUC of 0.91 (95% CI: 0.90-0.92) with a precision of 0.08 (95% CI: 0.08-0.08), and a recall of 0.73 (95% CI: 0.71-0.76). Performance was found to be robust across demographics, but cases with non-invasive cancers (AUC: 0.85, 95% CI: 0.83-0.87), calcifications (AUC: 0.80, 95% CI: 0.78-0.82), and dense breast tissue (AUC: 0.90, 95% CI: 0.88-0.91) were associated with significantly lower performance compared to other groups. These results highlight the need for detailed evaluation of model characteristics and vigilance in considering adoption of new tools for clinical deployment.

乳腺癌AI诊断模型评估医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。