用新架构提升多数据集医学图像分割精度,参数少三成却更准。
U-DFA: A Unified DINOv2-Unet with Dual Fusion Attention for Multi-Dataset Medical Segmentation
- 融合CNN空间特征与DINOv2语义特征,分阶段注入增强表示
- 在Synapse和ACDC数据集上达当前最优,参数仅需33%
- 适合追求高效高精度的医疗影像分割研究者
精准的医学图像分割对整体诊断至关重要,是诊断流程中最关键的任务之一。基于卷积神经网络(CNN)的模型虽广泛使用,但存在局部感受野局限,难以捕捉全局上下文信息。现有结合CNN与变换器的方法难以有效融合局部与全局特征。近年来,视觉语言模型(VLM)和基础模型被用于下游医学图像任务,但存在固有的领域差距和高计算开销问题。为此,我们提出U-DFA,一种统一的DINOv2-Unet编码器-解码器架构,引入新颖的局部-全局融合适配器(LGFA),以提升分割性能。LGFA模块将基于CNN的空域模式适配器(SPA)的空间特征注入到多个阶段的冻结DINOv2块中,实现高层语义与空间特征的有效融合。该方法在Synapse和ACDC数据集上达到当前最优表现,且仅需33%的可训练参数量。结果表明,U-DFA是一种在多种模态下鲁棒且可扩展的医学图像分割框架。
原文摘要 · Abstract (English)
Accurate medical image segmentation plays a crucial role in overall diagnosis and is one of the most essential tasks in the diagnostic pipeline. CNN-based models, despite their extensive use, suffer from a local receptive field and fail to capture the global context. A common approach that combines CNNs with transformers attempts to bridge this gap but fails to effectively fuse the local and global features. With the recent emergence of VLMs and foundation models, they have been adapted for downstream medical imaging tasks; however, they suffer from an inherent domain gap and high computational cost. To this end, we propose U-DFA, a unified DINOv2-Unet encoder-decoder architecture that integrates a novel Local-Global Fusion Adapter (LGFA) to enhance segmentation performance. LGFA modules inject spatial features from a CNN-based Spatial Pattern Adapter (SPA) module into frozen DINOv2 blocks at multiple stages, enabling effective fusion of high-level semantic and spatial features. Our method achieves state-of-the-art performance on the Synapse and ACDC datasets with only 33\% of the trainable model parameters. These results demonstrate that U-DFA is a robust and scalable framework for medical image segmentation across multiple modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。