arXiv:2508.06137eess.IVcs.CV2025-08被引 12

用Transformer+增强技术提升乳腺钼靶检测准确率与可解释性

Transformer-Based Explainable Deep Learning for Breast Cancer Detection in Mammography: The MammoFormer Framework

  • 结合Transformer与多特征增强,优化模型对局部和全局信息的捕捉
  • ViT达98.3%准确率,Swin Transformer在特定增强下提升13%
  • 提供多角度可解释性,适合临床医生信任与部署使用

乳腺癌钼靶检测因病灶微小且阅片者间差异大而困难。传统CNN难以兼顾局部细节与全局上下文,且缺乏可解释性,限制其临床应用。本文提出MammoFormer框架,融合Transformer架构、多特征增强组件与可解释AI功能。对比了七种模型(CNN、ViT、Swin Transformer、ConvNext)与四种增强方法(原始图像、负向变换、自适应直方图均衡化、方向梯度直方图)。通过针对架构的特征增强,实现最高13%性能提升;集成多视角可解释性,支持临床决策理解;构建结合CNN可靠性与Transformer全局建模能力的可部署集成系统。结果显示,适配增强后的Transformer可达到或超越CNN表现:在自适应直方图均衡化下ViT准确率达98.3%,在方向梯度直方图增强下Swin Transformer提升13.0%。

原文摘要 · Abstract (English)

Breast cancer detection through mammography interpretation remains difficult because of the minimal nature of abnormalities that experts need to identify alongside the variable interpretations between readers. The potential of CNNs for medical image analysis faces two limitations: they fail to process both local information and wide contextual data adequately, and do not provide explainable AI (XAI) operations that doctors need to accept them in clinics. The researcher developed the MammoFormer framework, which unites transformer-based architecture with multi-feature enhancement components and XAI functionalities within one framework. Seven different architectures consisting of CNNs, Vision Transformer, Swin Transformer, and ConvNext were tested alongside four enhancement techniques, including original images, negative transformation, adaptive histogram equalization, and histogram of oriented gradients. The MammoFormer framework addresses critical clinical adoption barriers of AI mammography systems through: (1) systematic optimization of transformer architectures via architecture-specific feature enhancement, achieving up to 13% performance improvement, (2) comprehensive explainable AI integration providing multi-perspective diagnostic interpretability, and (3) a clinically deployable ensemble system combining CNN reliability with transformer global context modeling. The combination of transformer models with suitable feature enhancements enables them to achieve equal or better results than CNN approaches. ViT achieves 98.3% accuracy alongside AHE while Swin Transformer gains a 13.0% advantage through HOG enhancements

乳腺癌检测Transformer可解释AI医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。