直接在k空间做MRI分类,效率提升68倍且更抗高加速干扰
Efficient Complex-Valued Vision Transformers for MRI Classification Directly from k-Space
- 设计新型复数视觉变压器kViT,用径向分块适配k空间频谱分布
- 在fastMRI和自建数据集上性能媲美主流图像模型,高加速下更稳定
- 训练时显存消耗减少68倍,适合临床扫描端直接部署
磁共振成像中的深度学习通常基于重建后的幅度图像,此过程丢失相位信息并需昂贵的变换。标准神经网络依赖局部操作(卷积或网格块),不适用于原始频域(k空间)数据的全局非局部特性。本文提出一种新型复数视觉变压器kViT,可直接在k空间执行分类。为弥合现有架构与磁共振物理之间的几何差异,引入径向k空间分块策略,尊重频域能量分布。在fastMRI和自建数据集上的大量实验表明,该方法分类性能与当前最优图像域基线(ResNet、EfficientNet、ViT)相当。关键优势在于对高加速因子具有更强鲁棒性,并实现计算效率范式转变:训练时显存消耗相比标准方法降低高达68倍。这为资源高效、直接从扫描仪获取数据的AI分析开辟了路径。
原文摘要 · Abstract (English)
Deep learning applications in Magnetic Resonance Imaging (MRI) predominantly operate on reconstructed magnitude images, a process that discards phase information and requires computationally expensive transforms. Standard neural network architectures rely on local operations (convolutions or grid-patches) that are ill-suited for the global, non-local nature of raw frequency-domain (k-Space) data. In this work, we propose a novel complex-valued Vision Transformer (kViT) designed to perform classification directly on k-Space data. To bridge the geometric disconnect between current architectures and MRI physics, we introduce a radial k-Space patching strategy that respects the spectral energy distribution of the frequency-domain. Extensive experiments on the fastMRI and in-house datasets demonstrate that our approach achieves classification performance competitive with state-of-the-art image-domain baselines (ResNet, EfficientNet, ViT). Crucially, kViT exhibits superior robustness to high acceleration factors and offers a paradigm shift in computational efficiency, reducing VRAM consumption during training by up to 68$\times$ compared to standard methods. This establishes a pathway for resource-efficient, direct-from-scanner AI analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。