arXiv:2504.11368cs.CV2025-04被引 21

用医生注视点+AI描述,提升弱监督医学图像分割精度

From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation

  • 融合医生注视点与视觉语言模型描述,互补提升分割能力
  • 在三个数据集上达80.78%~84.22%的Dice分数,比纯注视基线高3-5%
  • 适合追求可解释性与低标注成本的医疗AI研究者

医学图像分割因像素级标注成本高而困难。弱监督场景下,医生注视数据虽能捕捉诊断关注区域,但稀疏性限制其应用;而视觉语言模型(VLMs)提供语义信息却缺乏解释精度。本文提出教师-学生框架,融合两者优势:教师模型先利用增强的注视点与VLM生成的病灶形态描述学习,再通过多尺度特征对齐、置信度加权一致性约束和自适应掩码三种策略指导学生模型。在Kvasir-SEG、NCI-ISBI和ISIC数据集上,分割Dice分别达到80.78%、80.53%和84.22%,较仅使用注视点的基线提升3-5%,且不增加标注负担。该方法保持预测、注视与病灶描述间的关联性,保障临床可解释性。

原文摘要 · Abstract (English)

Medical image segmentation remains challenging due to the high cost of pixel-level annotations for training. In the context of weak supervision, clinician gaze data captures regions of diagnostic interest; however, its sparsity limits its use for segmentation. In contrast, vision-language models (VLMs) provide semantic context through textual descriptions but lack the explanation precision required. Recognizing that neither source alone suffices, we propose a teacher-student framework that integrates both gaze and language supervision, leveraging their complementary strengths. Our key insight is that gaze data indicates where clinicians focus during diagnosis, while VLMs explain why those regions are significant. To implement this, the teacher model first learns from gaze points enhanced by VLM-generated descriptions of lesion morphology, establishing a foundation for guiding the student model. The teacher then directs the student through three strategies: (1) Multi-scale feature alignment to fuse visual cues with textual semantics; (2) Confidence-weighted consistency constraints to focus on reliable predictions; (3) Adaptive masking to limit error propagation in uncertain areas. Experiments on the Kvasir-SEG, NCI-ISBI, and ISIC datasets show that our method achieves Dice scores of 80.78%, 80.53%, and 84.22%, respectively-improving 3-5% over gaze baselines without increasing the annotation burden. By preserving correlations among predictions, gaze data, and lesion descriptions, our framework also maintains clinical interpretability. This work illustrates how integrating human visual attention with AI-generated semantic context can effectively overcome the limitations of individual weak supervision signals, thereby advancing the development of deployable, annotation-efficient medical AI systems. Code is available at: https://github.com/jingkunchen/FGI.

医学图像分割弱监督视觉语言模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。