arXiv:2512.22185cs.CV2025-12

用双向注意力融合多模态数据,实现小动脉瘤的精准定位与可解释筛查。

Beyond Augmentation: Cross-Modal Transformer Fusion with Bi-directional Attention for Low-Data Aneurysm Screening

  • 通过双向注意力机制融合多模态影像,重构脑底动脉环解剖结构。
  • 在低样本下达到近完美的AUC-ROC,且对类别不平衡保持高精度。
  • 激活区域聚焦主血管,支持临床可解释的自动化筛查,适合医学影像分析者。

颅内动脉瘤破裂导致蛛网膜下腔出血,死亡率接近50%,早期检测至关重要。尽管CTA能快速筛查,但在复杂的脑底动脉环三维分支中检测微小动脉瘤仍依赖专家经验。现有自动化系统受限于类别不平衡、颅底伪影干扰、以及缺乏结构化定位的全局二分类,影响手术相关性与可解释性。本文提出CMTF-Net,一种跨模态目标融合框架,将动脉瘤筛查重构为解剖结构化的推理过程。通过独立监督14个血管区域,网络编码脑底动脉环几何特征,支持多段激活,契合临床工作流程。CMTF-Net在类别不平衡下实现近乎完美的AUC-ROC,置信区间窄,精度稳定。Grad-CAM与因果图显示激活区域沿主要动脉局部化,支持低数据场景下的可解释、解剖学基础筛查。

原文摘要 · Abstract (English)

Intracranial aneurysm rupture causes subarachnoid hemorrhage with mortality near 50%, making early detection critical. Although CTA enables rapid screening, detecting small aneurysms within the complex three-dimensional branching of the Circle of Willis remains expertise-dependent. Existing automated systems are constrained by class imbalance, skull-base artifacts that mimic vascular contrast, and reliance on global binary classification without structured localization, limiting surgical relevance and interpretability. We propose CMTF-Net, a cross-modal target fusion framework that reframes aneurysm screening as anatomically structured reasoning. By supervising 14 vascular territories independently, the network encodes Circle of Willis geometry while allowing multi-segment activation, aligning model design with clinical workflow. CMTF-Net achieves near-perfect AUC-ROC with narrow confidence intervals and sustained precision under imbalance. Grad-CAM and causal maps show spatially localized activation along major arteries, supporting interpretable, anatomically grounded screening in low-data settings.

医学影像跨模态融合可解释性小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。