arXiv:2608.09302cs.CV2026-08中稿 · Biomedical Signal …

用视觉语言模型提升宫腔镜手术场景分割精度,支持15类病灶精准定位。

Bootstrapping Vision-Language Model for Hysteroscopic Surgical Scene Segmentation

  • 基于预训练图像编码器与文本提示的视觉语言模型,实现像素级定位。
  • 在4020张高清图像上训练,分割性能显著优于现有方法。
  • 经多位妇科医生验证,模型泛化能力强,适合临床辅助决策场景。

宫腔镜手术场景分割对理解术中环境和实现计算机辅助干预至关重要。然而,由于不同病灶形态相似且存在反光、运动模糊、液体遮挡等伪影,该任务面临独特挑战。本文提出首个基于视觉语言模型(VLM)的宫腔镜手术场景分割方法,可对15类典型场景进行像素级定位。VLM-hyster采用预训练图像编码器提取鲁棒视觉特征,结合Transformer解码器实现密集预测,并设计类别专属文本提示,引入掩码蒸馏分支过滤低相关性特征,使模型更聚焦于类别特异性区域。我们构建了一个包含4020张高分辨率图像及详细标注掩码的大规模多中心数据集用于训练与评估。实验表明,VLM-hyster显著优于现有先进AI模型。多位妇科医生的评估及多中心前瞻性验证也证实其强鲁棒性与泛化能力。结果表明,VLM-hyster在实现宫腔镜手术中器械与病灶的AI辅助定位方面具有巨大潜力。代码已开源:https://github.com/viscom-tongji/VLM-hyster。

原文摘要 · Abstract (English)

Hysteroscopic surgical scene segmentation plays a pivotal role in understanding the hysteroscopic intraoperative environment as well as computer-assisted intervention. However, this task presents unique challenges due to the high morphological similarity among different lesions and the presence of artifacts such as specular reflections, motion blur, and fluid occlusions in surgical videos. In this work, we propose the first vision-language model (VLM)-based hysteroscopic surgical scene segmentation method, which performs pixel-wise localization for fifteen representative categories in hysteroscopic surgical scenes. Our VLM-hyster has a segmentation backbone that utilizes the pretrained image encoder for robust visual feature extraction, coupled with a transformer-based decoder for dense prediction. Moreover, we design category-specific text prompts and incorporate a masked distillation branch to filter out visual features with low correlation to the text prompts, enabling the model to focus more effectively on category-specific image regions and thereby enhancing segmentation performance. We collect a large multicentric hysteroscopic surgical scene dataset, containing 4,020 high-resolution images with detailed mask annotations, for model training and evaluation. Experimental results demonstrate that VLM-hyster substantially outperforms state-of-the-art AI models. Furthermore, extensive assessments by gynecologists, as well as multicentre and prospective validations, demonstrate VLM-hyster's robustness and generalizability. The results suggest that VLM-hyster earns considerable potential in enabling AI-assisted localization of surgical instruments and lesions in hysteroscopic surgeries. Code is available at https://github.com/viscom-tongji/VLM-hyster.

视觉语言模型手术分割宫腔镜医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。