arXiv:2509.11774cs.CV2025-09被引 2

轻量级模型提升视网膜血管分割精度与效率

SA-UNetv2: Rethinking Spatial Attention U-Net for Retinal Vessel Segmentation

  • 在所有跳跃连接中引入跨尺度空间注意力,强化多尺度特征融合
  • 采用加权BCE+MCC损失,在不平衡数据上表现更稳健,达最新精度
  • 参数仅0.26M,内存1.2MB,CPU推理快至1秒,适合边缘设备部署

视网膜血管分割对糖尿病视网膜病变、高血压及神经退行性疾病等早期诊断至关重要。尽管SA-UNet在瓶颈层引入空间注意力,但跳接路径中的注意力利用不足,且未解决严重前景-背景不平衡问题。我们提出SA-UNetv2,一种轻量级模型:将跨尺度空间注意力注入所有跳接路径,增强多尺度特征融合;并采用加权二值交叉熵(BCE)与马修斯相关系数(MCC)联合损失,提升对类别不平衡的鲁棒性。在公开的DRIVE和STARE数据集上,SA-UNetv2仅需1.2MB内存和0.26M参数(低于SA-UNet的50%),在592×592×3图像上实现1秒内CPU推理,达到当前最优性能,展现出强高效性与资源受限环境下的可部署性。

原文摘要 · Abstract (English)

Retinal vessel segmentation is essential for early diagnosis of diseases such as diabetic retinopathy, hypertension, and neurodegenerative disorders. Although SA-UNet introduces spatial attention in the bottleneck, it underuses attention in skip connections and does not address the severe foreground-background imbalance. We propose SA-UNetv2, a lightweight model that injects cross-scale spatial attention into all skip connections to strengthen multi-scale feature fusion and adopts a weighted Binary Cross-Entropy (BCE) plus Matthews Correlation Coefficient (MCC) loss to improve robustness to class imbalance. On the public DRIVE and STARE datasets, SA-UNetv2 achieves state-of-the-art performance with only 1.2MB memory and 0.26M parameters (less than 50% of SA-UNet), and 1 second CPU inference on 592 x 592 x 3 images, demonstrating strong efficiency and deployability in resource-constrained, CPU-only settings.

视网膜分割轻量模型注意力机制医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。