arXiv:2505.07349eess.IVcs.CV2025-05

用多平面视觉变换器提升不同方向MRI的出血分类准确率

Multi-Plane Vision Transformer for Hemorrhage Classification Using Axial and Sagittal MRI Data

  • 分路处理轴向与矢状面MRI,跨视角融合信息
  • AUC提升5.5%(对比ViT),优于传统CNN架构1.8%
  • 适合临床多角度MRI数据的出血检测任务

从磁共振成像(MRI)中识别脑出血对医疗人员至关重要。由于MRI扫描在对比度和方位上的多样性,使用神经网络识别出血存在挑战。传统方法常将图像重采样至固定平面,可能导致信息丢失。为此,我们提出一种用于处理不同方位数据的3D多平面视觉变换器(MP-ViT)。该模型为轴向与矢状面分别配置独立的Transformer编码器,并通过交叉注意力机制融合多视角信息。同时引入模态指示向量,补全缺失的对比度信息。在包含10,084例训练、1,289例验证和1,496例测试的临床真实数据集上,实验表明MP-ViT显著提升了曲线下面积(AUC),相比视觉变换器(ViT)提高5.5%,比基于CNN的架构高出1.8%。结果表明,当需结合多种方位对比时,MP-ViT具有显著提升出血检测性能的潜力。

原文摘要 · Abstract (English)

Identifying brain hemorrhages from magnetic resonance imaging (MRI) is a critical task for healthcare professionals. The diverse nature of MRI acquisitions with varying contrasts and orientation introduce complexity in identifying hemorrhage using neural networks. For acquisitions with varying orientations, traditional methods often involve resampling images to a fixed plane, which can lead to information loss. To address this, we propose a 3D multi-plane vision transformer (MP-ViT) for hemorrhage classification with varying orientation data. It employs two separate transformer encoders for axial and sagittal contrasts, using cross-attention to integrate information across orientations. MP-ViT also includes a modality indication vector to provide missing contrast information to the model. The effectiveness of the proposed model is demonstrated with extensive experiments on real world clinical dataset consists of 10,084 training, 1,289 validation and 1,496 test subjects. MP-ViT achieved substantial improvement in area under the curve (AUC), outperforming the vision transformer (ViT) by 5.5% and CNN-based architectures by 1.8%. These results highlight the potential of MP-ViT in improving performance for hemorrhage detection when different orientation contrasts are needed.

医学影像视觉变换器出血检测多视角融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。