arXiv:2506.10858eess.IVcs.CV2025-06被引 1

纯VRWKV模型在医学图像分割中表现优异,性能超越现有方法。

Med-URWKV†: Toward Enhanced Pretrained Pure VRWKV Models for Medical Image Segmentation

  • 基于预训练纯VRWKV架构,构建轻量级分割模型
  • 引入波频注意力与多尺度通道融合模块,提升细节捕捉能力
  • 参数减半仍达88.00%最高Dice,适合医疗视觉任务

医学图像分割是辅助诊疗的基础任务。现有基于CNN、ViT、Mamba及混合模型的方法仍受限于感受野有限、计算成本高或精度不足。近期视觉感受野加权键值(VRWKV)模型展现出建模长程依赖的潜力。然而,当前基于VRWKV的医学图像分割研究多集中于从零开始训练的混合架构,大规模预训练纯VRWKV模型的潜力尚未探索。本文系统评估纯VRWKV架构在该任务中的有效性,通过复用不同规模的预训练VRWKV编码器并搭配纯VRWKV解码器,构建了Med-URWKV-T和Med-URWKV-S,实现对预训练纯VRWKV模型的全面评估。为进一步提升性能,提出两个兼容VRWKV的模块:频率感知小波注意力(FAWA),利用小波变换捕捉边缘细节与结构特征;多尺度通道融合(MSCF),整合多尺度特征以增强信息通道表征。将二者融入Med-URWKV-T,得到增强模型Med-URWKV†。在五个医学图像分割数据集上的实验表明,Med-URWKV性能达到或优于当前最先进方法及精心设计的混合VRWKV架构。其中,Med-URWKV†在参数量仅为Med-URWKV-S一半的情况下,分割准确率更高,平均Dice相似系数达88.00%,为最高水平。代码将公开。

原文摘要 · Abstract (English)

Medical image segmentation is a fundamental task in computer-aided diagnosis and treatment. Existing approaches based on CNNs, ViTs, Mamba, and hybrid models still suffer from limitations such as restricted receptive fields, high computational cost, or insufficient accuracy. Recently, Vision Receptive-field Weighted Key-Value (VRWKV) models have emerged as a promising alternative,delivering strong long-range dependency modeling for visual tasks. However, current studies on VRWKV-based medical image segmentation mainly focus on hybrid architectures trained from scratch, while the potential of large-scale pretrained pure VRWKV models remains unexplored. In this work, we systematically investigate the effectiveness of pure VRWKV architectures for medical image segmentation. We construct Med-URWKV-T and Med-URWKV-S by reusing pretrained VRWKV encoders at different scales and pairing them with pure VRWKV decoders, enabling a comprehensive evaluation of pretrained pure VRWKV models in this domain. To further enhance performance, we propose two VRWKV-compatible modules: a Frequency-Aware Wavelet Attention (FAWA) module, which exploits wavelet transforms to capture edge details and structural characteristics, and a Multi-Scale Channel Fusion (MSCF) module, which integrates multi-scale features to strengthen informative channel representations. By incorporating them into Med-URWKV-T, we obtain the enhanced model Med-URWKV†. Extensive experiments on five medical image segmentation datasets demonstrate that Med-URWKV achieves performance comparable to or superior to state-of-the-art methods and carefully designed hybrid VRWKV architectures. Moreover, Med-URWKV† further improves segmentation accuracy, surpassing Med-URWKV-S while using only half of its parameter count, and achieves the highest average Dice similarity coefficient of 88.00%. The codes will be released.

医学图像分割VRWKV小波注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。