arXiv:2607.06309cs.CVcs.AI2026-07

用令牌级交互提升乳腺癌分类多视角融合效果

Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification

论文配图:Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification
图 1 · 摘自论文原文
  • 设计令牌中心的双视角融合框架,通过专用融合令牌实现跨视角通信
  • 在多个Transformer深度插入融合模块,实现多层次信息交互,提升分类准确率
  • 适用于医学影像多视角分析,尤其适合需要精细特征融合的乳腺癌诊断

从乳腺钼靶图像中准确分类乳腺癌需有效整合头尾向(CC)和内外斜向(MLO)视图提供的互补信息,以更完整刻画乳腺异常。现有方法通常依赖特征级聚合或单阶段交叉注意力,易混淆视图特异与共享表征,且交互受限于有限网络深度。为此,我们提出一种以令牌为中心的双视角学习框架,将提示调优与跨视图融合统一于冻结的视觉Transformer骨干网络中。该框架将视间交互重构为结构化的令牌级通信:专用融合令牌通过交叉注意力显式编码CC与MLO视图间的双向信息交换,作为跨视图依赖关系的中间载体,而非直接特征融合。不同于传统单层融合,融合模块被插入至多个Transformer深度,实现编码器层次上的渐进式、重复性交互。融合令牌重新融入令牌序列并由后续层优化,促进互补信息的层级传播,同时保持视图特异性结构。在VinDr-Mammo与CMMD数据集上的实验表明,该方法显著优于线性探测、仅提示调优及传统融合基线。在VinDr-Mammo BI-RADS分类任务中,取得50.40%的F1分数与0.8090的AUC,较双视角融合基线在二分类设置下提升0.10 AUC。消融实验证实了基于令牌的融合与多深度交互设计的有效性。

原文摘要 · Abstract (English)

Accurate breast cancer classification from mammography requires effective integration of complementary information from craniocaudal (CC) and mediolateral oblique (MLO) views, which provide a more complete characterization of breast abnormalities. However, existing multi-view learning approaches typically rely on feature-level aggregation or single-stage cross-attention, which can entangle view-specific and shared representations and restrict interaction to limited network depths. To address these limitations, we propose a token-centric dual-view learning framework that unifies prompt-based adaptation and cross-view fusion within a frozen vision transformer backbone. The framework reformulates inter-view interaction as structured token-level communication, where dedicated fusion tokens explicitly encode bidirectional information exchange between CC and MLO views via cross-attention, serving as intermediate carriers of cross-view dependencies rather than relying on direct feature fusion. Unlike conventional methods that apply fusion at a single layer, fusion modules are inserted at multiple transformer depths, enabling progressive and repeated interaction across the encoder hierarchy. Fusion tokens are reintegrated into the token sequence and refined by subsequent transformer layers, facilitating hierarchical propagation of complementary information while preserving view-specific structure. Experiments on VinDr-Mammo and CMMD datasets demonstrate consistent improvements over linear probing, prompt-only adaptation, and conventional fusion baselines. On the VinDr-Mammo BI-RADS classification task, the framework achieves 50.40% F1-score and 0.8090 AUC, including a 0.10 AUC improvement over a dual-view fusion baseline in the binary setting. Ablation studies further validate the effectiveness of token-based fusion and multi-depth interaction design.

医学影像多视图学习视觉变压器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。