用多窗口注意力机制提升3D MRI预测胆管癌神经侵犯的精度
MMA-Former: Multi-Window Mixture-of-Head Attention Transformer for Adaptive PNI Prediction in 3D MRI

- 采用并行多尺度结构与窗口自适应注意力机制
- 在168例数据上达AUC 0.752,优于现有模型
- 适合需要精准影像分析的临床研究者
神经侵犯(PNI)是胆管癌的重要预后因子。基于3D MRI进行无创预测极具挑战,要求模型能高效捕捉细微结构与全局上下文。本文提出多窗口混合头注意力变换器(MMA-Former),一种新型端到端3D架构,包含粗细特征并行提取的粗细变换器(CFT)结构。通过引入创新的窗口特定混合头注意力(WS-MoH)机制,该模型为每个3D窗口生成独立表征,并动态将窗口路由至专用或共享注意力头,实现空间自适应特征提取,在不增加参数量的前提下提升特征专属性并降低冗余。在包含168例T1加权MRI扫描的回顾性数据集上,MMA-Former取得AUC 0.752,优于最佳CNN(AUC 0.708)和基准Transformer(AUC 0.681)。
原文摘要 · Abstract (English)
Perineural invasion (PNI) is a critical prognostic factor in cholangiocarcinoma. Non-invasive prediction from 3D MRI is challenging, demanding models that efficiently capture both fine-grained details and global context. We propose the Multi-window Mixture-of-Head Attention Transformer (MMA-Former), a novel end-to-end 3D architecture featuring a Coarse-Fine Transformer (CFT) structure for parallel multi-scale feature extraction. We advance this structure by integrating a novel Window-Specific Mixture-of-Head attention (WS-MoH) mechanism. Unlike standard Multi-Head Self Attention (MSA), WS-MoH generates a representation for each 3D window and dynamically routes the entire window to specialized or common attention heads. This enables spatially adaptive feature extraction tailored to the local context of each window, enhancing specialization and reducing redundancy without increasing parameters. Evaluated on a retrospective dataset of 168 T1-weighted MRI scans, MMA-Former achieved an AUC of 0.752, outperforming other 3D architectures, including the best CNN (AUC of 0.708) and Transformer baselines (AUC of 0.681).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。