arXiv:2511.17828cs.CVcs.AI2025-11被引 1

用基础模型提升乳腺影像密度分类,支持多模态数据且可解释。

Toward explainable AI approaches for breast imaging: adapting foundation models to diverse populations

  • 基于BiomedCLIP构建多模态乳腺影像分类模型,解决数据不平衡问题。
  • 多模态模型准确率74%,跨数据集AUC达0.80–0.93,优于单模态。
  • 可视化验证注意力模式临床合理,适合医疗决策辅助场景。

基础模型在医学影像中潜力巨大,但其在乳腺影像中的应用仍不充分。本研究采用BiomedCLIP作为基础模型,利用多模态乳腺钼靶数据(合成2D图像、数字乳腺摄影、数字乳腺断层成像)实现自动化BI-RADS乳腺密度分类。基于96,995张图像,比较了仅使用2D图像的单模态与多模态训练方法,通过加权对比学习缓解类别不平衡。两种方法准确率相近(多模态:0.74,单模态:0.73),但多模态模型在不同成像模态下更具泛化能力,各BI-RADS类别AUC均高于0.84。在RSNA和EMBED外部数据集上验证,模型表现稳定,AUC范围为0.80–0.93。GradCAM可视化显示一致且符合临床意义的关注区域,证明模型具备可解释性与鲁棒性。研究证实基础模型在乳腺影像任务中的可行性,为未来诊断应用提供新路径。

原文摘要 · Abstract (English)

Foundation models hold promise for specialized medical imaging tasks, though their effectiveness in breast imaging remains underexplored. This study leverages BiomedCLIP as a foundation model to address challenges in model generalization. BiomedCLIP was adapted for automated BI-RADS breast density classification using multi-modality mammographic data (synthesized 2D images, digital mammography, and digital breast tomosynthesis). Using 96,995 images, we compared single-modality (s2D only) and multi-modality training approaches, addressing class imbalance through weighted contrastive learning. Both approaches achieved similar accuracy (multi-modality: 0.74, single-modality: 0.73), with the multi-modality model offering broader applicability across different imaging modalities and higher AUC values consistently above 0.84 across BI-RADS categories. External validation on the RSNA and EMBED datasets showed strong generalization capabilities (AUC range: 0.80-0.93). GradCAM visualizations confirmed consistent and clinically relevant attention patterns, highlighting the models interpretability and robustness. This research underscores the potential of foundation models for breast imaging applications, paving the way for future extensions for diagnostic tasks.

乳腺影像基础模型可解释性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。