arXiv:2607.04249cs.CV2026-07中稿 · ECCV

针对医学影像分割中少量标注数据的偏差问题,提出分布感知采样与记忆引导增强新方法。

Beyond Random Sampling: Distribution-Aware Alignment for Semi-Supervised Medical Image Segmentation

论文配图:Beyond Random Sampling: Distribution-Aware Alignment for Semi-Supervised Medical Image Segmentation
图 1 · 摘自论文原文
  • 用视觉基础模型和密度聚类选代表性样本,构建更均衡的标签数据集。
  • 在1/16低标注比例下,分割精度显著优于随机采样方法。
  • 适合标注稀缺的医学图像分析场景,尤其对类别不平衡数据有效。

精准的医学图像分割对临床诊断和治疗规划至关重要,但依赖昂贵的专业标注。半监督医学图像分割(SSMIS)虽能降低成本,但通常假设数据独立同分布,采用随机采样。这种策略在大数据量下统计有效,但在小样本情况下会因表示偏差而无法捕捉医学数据的异质结构。为此,我们提出一种以分布对齐为核心的高数据效率框架。首先,引入离线分布感知样本选择策略,利用视觉基础模型(VFMs)和自研的Density-K-Center算法,显式识别出具有代表性的结构锚点,构建更具代表性的标注域。其次,为弥合剩余分布差异,提出记忆引导的复制粘贴(MCP)模块。该模块针对医学影像固有的类别不平衡问题,通过语义记忆机制检索历史一致先验,实现跨域对齐,促进语义一致性。结合易到难渐进式训练策略,有效缓解初期伪标签噪声。在六个多样化的2D和3D数据集上进行的大量实验表明,该框架在极低标注比例(如1/16)下仍具备强大分割性能。

原文摘要 · Abstract (English)

Precise medical image segmentation is crucial for clinical diagnosis and treatment planning, yet relies heavily on expensive expert annotations. Semi-supervised medical image segmentation (SSMIS) offers a cost-effective solution but typically operates under the assumption of independent and identically distributed (i.i.d.) data, defaulting to random sampling. While statistically valid at scale, this strategy suffers from severe representation bias in low-data regimes, failing to capture the heterogeneous medical data manifold. To address this, we propose a highly data-efficient framework driven by distribution alignment. First, we introduce an offline Distribution-Aware Sample Selection strategy. By leveraging Vision Foundation Models (VFMs) and our designed Density-K-Center algorithm, we explicitly identify representative structural anchors, establishing a more representative labeled domain. Second, to bridge the remaining distribution gap, we propose the Memory-guided Copy-Paste (MCP) module. Tailored for the inherent class imbalance in medical scans, MCP leverages a semantic memory mechanism to retrieve historically consistent priors for cross-domain alignment, encouraging semantic consistency. Coupled with an easy-to-hard progressive schedule, this framework effectively mitigates early-stage pseudo-label noise. Extensive experiments on six diverse 2D and 3D datasets demonstrate strong segmentation performance, particularly in extremely low-labeled scenarios (\eg, 1/16 ratio).

医学图像半监督分布对齐低样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。