arXiv:2604.11164cs.CV2026-04

用细粒度视觉特征提升少标注医学图像分割精度

RADA: Region-Aware Dual-encoder Auxiliary learning for Barely-supervised Medical Image Segmentation

论文配图:RADA: Region-Aware Dual-encoder Auxiliary learning for Barely-supervised Medical Image Segmentation
图 1 · 摘自论文原文
  • 双编码器结构结合图像与文本信息,捕捉区域特异性特征
  • 在仅标注少量切片下,三个数据集均达当前最优性能
  • 适合标注稀缺的医学图像分割场景,尤其3D体积数据

深度学习极大推动了医学图像分割发展,但其成功依赖于全监督学习,需对3D体扫描进行密集标注,成本高昂。少标注学习通过每体积仅标注少数切片降低标注负担。现有方法通常利用几何连续性将稀疏标注传播至未标注切片生成伪标签,但缺乏语义理解,常导致伪标签质量低。医学图像分割本质是像素级视觉理解任务,精度根本取决于局部细粒度视觉特征质量。受此启发,我们提出RADA:一种基于Alpha-CLIP预训练的区域感知双编码器辅助学习框架,从原始图像和有限标注中提取细粒度、区域特异性视觉特征。该框架融合图像级细粒度特征与文本级语义引导,提供区域感知的语义监督,连接图像级语义与像素级分割。集成于三视图训练框架,在LA2018、KiTS19和LiTS数据集上实现极稀疏标注下的最先进性能,展现出跨数据集的强泛化能力。

原文摘要 · Abstract (English)

Deep learning has greatly advanced medical image segmentation, but its success relies heavily on fully supervised learning, which requires dense annotations that are costly and time-consuming for 3D volumetric scans. Barely-supervised learning reduces annotation burden by using only a few labeled slices per volume. Existing methods typically propagate sparse annotations to unlabeled slices through geometric continuity to generate pseudo-labels, but this strategy lacks semantic understanding, often resulting in low-quality pseudo-labels. Furthermore, medical image segmentation is inherently a pixel-level visual understanding task, where accuracy fundamentally depends on the quality of local, fine-grained visual features. Inspired by this, we propose RADA, a novel Region-Aware Dual-encoder Auxiliary learning pipeline which introduces a dual-encoder framework pre-trained on Alpha-CLIP to extract fine-grained, region-specific visual features from the original images and limited annotations. The framework combines image-level fine-grained visual features with text-level semantic guidance, providing region-aware semantic supervision that bridges image-level semantics and pixel-level segmentation. Integrated into a triple-view training framework, RADA achieves SOTA performance under extremely sparse annotation settings on LA2018, KiTS19 and LiTS, demonstrating robust generalization across diverse datasets.

医学图像分割少标注学习双编码器视觉特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。