arXiv:2503.13309eess.IVcs.AI2025-03被引 4

用多尺度多视角Transformer提升乳腺癌诊断,更抗缺失视图。

Integrating AI for Human-Centric Breast Cancer Diagnostics: A Multi-Scale and Multi-View Swin Transformer Framework

  • 结合SAM分割乳腺组织,多尺度捕捉肿瘤局部与周围结构特征。
  • 融合上下文与局部信息,输出更符合医生阅片习惯。
  • 设计混合融合结构,单视图也能保持诊断鲁棒性。

尽管计算机辅助诊断(CAD)系统取得进展,乳腺癌仍是全球女性癌症致死的主要原因。人工智能(AI)在深度学习架构方面展现出显著潜力,可用于乳腺钼靶诊断。然而,现有方法常忽视对详细肿瘤标注的依赖及测试时缺失视图的问题。为此,本文提出一种基于Swin Transformer的多尺度、多视图混合框架(MSMV-Swin),增强诊断的鲁棒性与准确性。该框架作为辅助工具,帮助放射科医生更高效分析多视图钼靶图像。具体而言,利用分割任意模型(SAM)分离乳腺腺体区域,降低背景噪声,实现全面特征提取。多尺度设计兼顾肿瘤区域与周围组织的空间特性,同时捕捉局部与上下文信息。融合局部与上下文数据使输出更贴近医生读片逻辑,促进人机协作与信任。此外,设计了混合融合结构以应对临床中常见仅存单一视图的情况,提升实际应用适应性。

原文摘要 · Abstract (English)

Despite advancements in Computer-Aided Diagnosis (CAD) systems, breast cancer remains one of the leading causes of cancer-related deaths among women worldwide. Recent breakthroughs in Artificial Intelligence (AI) have shown significant promise in development of advanced Deep Learning (DL) architectures for breast cancer diagnosis through mammography. In this context, the paper focuses on the integration of AI within a Human-Centric workflow to enhance breast cancer diagnostics. Key challenges are, however, largely overlooked such as reliance on detailed tumor annotations and susceptibility to missing views, particularly during test time. To address these issues, we propose a hybrid, multi-scale and multi-view Swin Transformer-based framework (MSMV-Swin) that enhances diagnostic robustness and accuracy. The proposed MSMV-Swin framework is designed to work as a decision-support tool, helping radiologists analyze multi-view mammograms more effectively. More specifically, the MSMV-Swin framework leverages the Segment Anything Model (SAM) to isolate the breast lobe, reducing background noise and enabling comprehensive feature extraction. The multi-scale nature of the proposed MSMV-Swin framework accounts for tumor-specific regions as well as the spatial characteristics of tissues surrounding the tumor, capturing both localized and contextual information. The integration of contextual and localized data ensures that MSMV-Swin's outputs align with the way radiologists interpret mammograms, fostering better human-AI interaction and trust. A hybrid fusion structure is then designed to ensure robustness against missing views, a common occurrence in clinical practice when only a single mammogram view is available.

乳腺癌多视图Swin TransformerAI辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。