arXiv:2604.10754eess.IV2026-04被引 1

用人类注视数据提升医学图像分割的标注效率与精度。

Human Gaze-based Dual Teacher Guidance Learning for Semi-Supervised Medical Image Segmentation

  • 引入注视数据作为双重教师指导,增强模型感知能力。
  • 在十个不同器官/组织上均优于现有方法,泛化性强。
  • 适合医疗图像标注资源稀缺场景下的半监督学习研究者。

在医学图像分割领域,标注数据稀缺是制约模型精准定位目标区域的主要挑战。相比人工标注,注视数据更易获取且成本更低。基于经典的半监督学习框架Mean-Teacher,本文提出人眼注视引导的双教师指导学习模型(HG-DTGL)。通过融合注视数据,解决两个关键问题:1)在有限标注数据下扩展数据集规模与多样性;2)增强网络感知能力。设计了GazeMix生成可靠混合数据以扩充数据多样性,引入多尺度注视感知(MGP)模块提取网络多尺度感知特征,并构建注视损失(Gaze Loss)使模型关注与人眼注视对齐。在多个不同模态的数据集上验证,覆盖十种不同器官/组织,实验结果表明本方法在多种场景下表现优异,具备强泛化能力,充分展示了注视数据在半监督医学图像分割中的巨大应用潜力。

原文摘要 · Abstract (English)

In the field of medical image segmentation, the scarcity of labeled data poses a major challenge for existing models to accurately perceive target regions. Compared with manual annotation, gaze data is easier and cheaper to obtain. As a classical semi-supervised learning framework, mean-teacher can effectively use a large number of unlabeled medical images for stable training through self-teaching and collaborative optimization. Our study is based on the mean-teacher framework. By combining gaze data, it aims to address two crucial issues in semi-supervised medical image segmentation: 1) expand the scale and diversity of the dataset with limited labeled data; 2) enhance the network's perception ability. We propose the Human Gaze-based Dual Teacher Guidance Learning model (HG-DTGL). In this model, human gaze serves as an additional hidden `teacher' in the mean-teacher architecture. We introduce the GazeMix to generate reliable mixed data to expand the diversity and scale of the dataset, and the Multi-scale Gaze Perception (MGP) module is used to extract the multi-scale perception of the network. A Gaze Loss is designed to align the model's perception with human gaze. We have verified HG-DTGL on multiple datasets of different modalities and achieved superior performance on a total of ten different organs/tissues, with extensive experiments. This demonstrates that our method has strong generalization ability for medical images of different modalities, and shows the great application potential of gaze data in semi-supervised medical image segmentation.

医学图像半监督学习注视数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。