通过聚类发现乳腺癌风险模型中的重复影像特征。
Revealing Mammographic Phenotypes in Deep Learning Breast Cancer Risk Models
- 用预训练模型嵌入聚类,识别跨人群的重复影像模式。
- 高风险特征与致密组织、微钙化等结构相关,5年风险升高。
- 揭示了模型潜在混淆因子,适合临床与AI交叉研究者。
基于乳腺钼靶的深度学习模型提升了乳腺癌风险预测能力,但其学习到的影像模式仍缺乏深入探索。现有可解释性方法依赖单张图像的显著性图,无法识别大规模患者队列中反复出现的哺乳影像表型。通过聚类预训练模型Mirai的图像块嵌入,我们分离出与5年癌症风险相关的重复表型。分析显示,增加风险的表型捕捉了复杂结构(如致密组织、微钙化)和捷径伪影(如夹子)。这些表型与年龄较大及更高BI-RADS密度显著相关。该框架将组织模式与AI风险评分关联,揭示了临床特征及潜在的模型混淆因素。
原文摘要 · Abstract (English)
Mammogram-based deep learning models have improved breast cancer risk prediction, but the learned imaging patterns remain underexplored. Existing interpretability methods rely on single-image saliency maps, failing to identify recurring mammographic phenotypes across large patient cohorts. By clustering patch embeddings from a pre-trained model, Mirai, we isolate recurring phenotypes linked to 5-year cancer risk. Analyses show risk-increasing phenotypes capture complex structures (e.g., dense tissue, microcalcifications) and shortcut artifacts (e.g., clips). These phenotypes correlate strongly with older age and higher BI-RADS density. Our framework connects tissue patterns to AI risk scores, revealing clinical signatures and potential latent model confounders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。