用跨模态自适应融合技术,让不配对的CT/MRI数据联合训练肝肿瘤分割模型。
A-QCF-Net: An Adaptive Quaternion Cross-Fusion Network for Multimodal Liver Tumor Segmentation from Unpaired Datasets
- 通过四元数网络构建共享特征空间,实现CT与MRI间双向知识迁移。
- 在未配对的LiTS和ATLAS数据上,肿瘤分割Dice分别达76.7%和78.3%。
- 适合缺乏配对影像数据的医院或研究者,推动医疗影像数据价值释放。
多模态医学影像能提供互补信息,对病灶精准分割至关重要,但深度学习模型的发展受限于不同模态配对且空间对齐的大规模数据集稀缺。本文提出自适应四元数交叉融合网络(A-QCF-Net),从完全分离的CT和MRI队列中学习单一统一的分割模型。该架构利用四元数神经网络的参数效率与表达能力构建共享特征空间。核心为自适应四元数交叉融合(A-QCF)模块,一种数据驱动的注意力机制,实现双流间双向知识传递。通过动态调节信息流,A-QCF模块使网络交换抽象的模态特异性知识,如CT中的清晰解剖边界与MRI中的细微软组织对比。这种相互交流正则化并丰富了双流特征表示。我们在未配对的LiTS(CT)和ATLAS(MRI)数据集上联合训练单个模型,所得模型在CT上的肿瘤Dice达到76.7%,在MRI上达78.3%,显著优于强基线unimodal nnU-Net,提升分别为5.4%和4.7%。进一步的Grad-CAM与Grad-CAM++可解释性分析表明,模型正确聚焦于相关病灶结构,确保学习到的表征具有临床意义。这为挖掘医疗中常见的大量未配对影像资源提供了稳健且临床可行的新范式。
原文摘要 · Abstract (English)
Multimodal medical imaging provides complementary information that is crucial for accurate delineation of pathology, but the development of deep learning models is limited by the scarcity of large datasets in which different modalities are paired and spatially aligned. This paper addresses this fundamental limitation by proposing an Adaptive Quaternion Cross-Fusion Network (A-QCF-Net) that learns a single unified segmentation model from completely separate and unpaired CT and MRI cohorts. The architecture exploits the parameter efficiency and expressive power of Quaternion Neural Networks to construct a shared feature space. At its core is the Adaptive Quaternion Cross-Fusion (A-QCF) block, a data driven attention module that enables bidirectional knowledge transfer between the two streams. By learning to modulate the flow of information dynamically, the A-QCF block allows the network to exchange abstract modality specific expertise, such as the sharp anatomical boundary information available in CT and the subtle soft tissue contrast provided by MRI. This mutual exchange regularizes and enriches the feature representations of both streams. We validate the framework by jointly training a single model on the unpaired LiTS (CT) and ATLAS (MRI) datasets. The jointly trained model achieves Tumor Dice scores of 76.7% on CT and 78.3% on MRI, significantly exceeding the strong unimodal nnU-Net baseline by margins of 5.4% and 4.7% respectively. Furthermore, comprehensive explainability analysis using Grad-CAM and Grad-CAM++ confirms that the model correctly focuses on relevant pathological structures, ensuring the learned representations are clinically meaningful. This provides a robust and clinically viable paradigm for unlocking the large unpaired imaging archives that are common in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。