arXiv:2602.07819eess.IVcs.CV2026-02

用外部大模型指导医学图像分割,解决少数类别识别难题

DINO-Mix: Distilling Foundational Knowledge with Cross-Domain CutMix for Semi-supervised Class-imbalanced Medical Image Segmentation

  • 引入DINOv3作为无偏外部教师,提供跨领域语义指导
  • 在Synapse和AMOS数据集上显著提升少数类分割性能
  • 适合医疗图像少样本、类别不平衡场景的研究者

半监督学习(SSL)已成为缓解医学图像分割密集标注成本的关键范式。然而,现有SSL框架本质上是‘向内看’,仅在目标数据集内部循环信息与偏差,导致在类别不平衡下陷入确认偏差的恶性循环,致使少数类别识别完全失败。为打破这一系统性问题,我们提出一种多层级‘向外看’的新范式。核心创新是基础知识蒸馏(FKD),通过引入预训练视觉基础模型DINOv3作为无偏外部语义教师,超越医学影像范畴,提供稳定的跨域监督信号,锚定少数类的学习。为进一步拓展外部视角,我们提出渐进式不平衡感知CutMix(PIC),构建动态课程,自适应地迫使模型在有标签和无标签子集中关注少数类。该分层策略形成我们的DINO-Mix框架,在挑战性的半监督类别不平衡医学图像分割基准Synapse和AMOS上取得显著性能提升。

原文摘要 · Abstract (English)

Semi-supervised learning (SSL) has emerged as a critical paradigm for medical image segmentation, mitigating the immense cost of dense annotations. However, prevailing SSL frameworks are fundamentally "inward-looking", recycling information and biases solely from within the target dataset. This design triggers a vicious cycle of confirmation bias under class imbalance, leading to the catastrophic failure to recognize minority classes. To dismantle this systemic issue, we propose a paradigm shift to a multi-level "outward-looking" framework. Our primary innovation is Foundational Knowledge Distillation (FKD), which looks outward beyond the confines of medical imaging by introducing a pre-trained visual foundation model, DINOv3, as an unbiased external semantic teacher. Instead of trusting the student's biased high confidence, our method distills knowledge from DINOv3's robust understanding of high semantic uniqueness, providing a stable, cross-domain supervisory signal that anchors the learning of minority classes. To complement this core strategy, we further look outward within the data by proposing Progressive Imbalance-aware CutMix (PIC), which creates a dynamic curriculum that adaptively forces the model to focus on minority classes in both labeled and unlabeled subsets. This layered strategy forms our framework, DINO-Mix, which breaks the vicious cycle of bias and achieves remarkable performance on challenging semi-supervised class-imbalanced medical image segmentation benchmarks Synapse and AMOS.

医学图像半监督学习类别不平衡知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。