融合多视图与多任务的深度模型,提升乳腺钼靶诊断准确率。
A Hybrid CNN-VSSM model for Multi-View, Multi-Task Mammography Analysis: Robust Diagnosis with Attention-Based Fusion
- 用混合CNN-VSSM结构同时处理四视角影像,捕捉局部与全局特征。
- 在二分类任务中达到0.9967 AUC和0.9830 F1,三分类任务F1达0.7790。
- 动态注意力融合机制有效应对缺失数据,适合临床辅助诊断场景。
早期精准解读筛查乳腺钼靶片对乳腺癌的有效检测至关重要,但因影像表现细微且诊断存在模糊性,仍具挑战。现有多数AI方法仅关注单视图输入或单任务输出,限制了其临床应用。为此,我们提出一种新型多视图、多任务混合深度学习框架,可同时处理标准四视角乳腺钼靶片,并联合预测每侧乳房的诊断标签与BI-RADS评分。该架构结合卷积编码器提取丰富局部特征,以及视觉状态空间模型(VSSMs)捕获全局上下文依赖关系。为提升鲁棒性与可解释性,引入门控注意力融合模块,动态加权各视图信息,有效处理缺失数据情况。我们在不同复杂度的任务上开展广泛实验,对比所提混合模型与基线CNN及VSSM模型在单任务与多任务学习中的表现。结果表明,混合模型始终优于基线:在二分类(BI-RADS 1 vs. 5)任务中,共享模型达到0.9967 AUC与0.9830 F1;三分类任务中F1为0.7790;五分类任务最佳F1达0.4904。这些结果验证了所提框架的有效性,也揭示了多任务学习在提升诊断性能方面的潜力与局限。
原文摘要 · Abstract (English)
Early and accurate interpretation of screening mammograms is essential for effective breast cancer detection, yet it remains a complex challenge due to subtle imaging findings and diagnostic ambiguity. Many existing AI approaches fall short by focusing on single view inputs or single-task outputs, limiting their clinical utility. To address these limitations, we propose a novel multi-view, multitask hybrid deep learning framework that processes all four standard mammography views and jointly predicts diagnostic labels and BI-RADS scores for each breast. Our architecture integrates a hybrid CNN VSSM backbone, combining convolutional encoders for rich local feature extraction with Visual State Space Models (VSSMs) to capture global contextual dependencies. To improve robustness and interpretability, we incorporate a gated attention-based fusion module that dynamically weights information across views, effectively handling cases with missing data. We conduct extensive experiments across diagnostic tasks of varying complexity, benchmarking our proposed hybrid models against baseline CNN architectures and VSSM models in both single task and multi task learning settings. Across all tasks, the hybrid models consistently outperform the baselines. In the binary BI-RADS 1 vs. 5 classification task, the shared hybrid model achieves an AUC of 0.9967 and an F1 score of 0.9830. For the more challenging ternary classification, it attains an F1 score of 0.7790, while in the five-class BI-RADS task, the best F1 score reaches 0.4904. These results highlight the effectiveness of the proposed hybrid framework and underscore both the potential and limitations of multitask learning for improving diagnostic performance and enabling clinically meaningful mammography analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。