arXiv:2507.04317eess.IVcs.AI2025-07被引 1

用对比学习与强化学习提升腹腔镜手术场景分割精度

CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning

  • 结合对比语言-视觉预训练与强化学习,动态优化分割结果
  • 在EndoVis 2018上达到81%的平均IoU,优于现有模型
  • 适合需要高精度手术图像理解的临床辅助系统

理解手术场景可提升患者医疗质量,尤其在微创手术(MIS)中生成大量视频数据的背景下。本文提出CLIP-RL,一种专用于手术场景语义分割的对比语言-图像预训练模型。该方法融合强化学习与课程学习,实现训练全程对分割掩码的持续优化。模型在不同光学条件下表现稳健,包括遮挡、纹理变化和动态光照等挑战。CLIP模型作为强大特征提取器,捕捉丰富语义上下文,增强器械与组织的区分能力;强化学习模块通过迭代动作空间调整,动态优化预测。在EndoVis 2018和EndoVis 2017数据集上评估,CLIP-RL分别取得81%和74.12%的平均交并比(mean IoU),显著优于当前最优模型,性能提升源于对比学习、强化学习与课程学习的协同作用。

原文摘要 · Abstract (English)

Understanding surgical scenes can provide better healthcare quality for patients, especially with the vast amount of video data that is generated during MIS. Processing these videos generates valuable assets for training sophisticated models. In this paper, we introduce CLIP-RL, a novel contrastive language-image pre-training model tailored for semantic segmentation for surgical scenes. CLIP-RL presents a new segmentation approach which involves reinforcement learning and curriculum learning, enabling continuous refinement of the segmentation masks during the full training pipeline. Our model has shown robust performance in different optical settings, such as occlusions, texture variations, and dynamic lighting, presenting significant challenges. CLIP model serves as a powerful feature extractor, capturing rich semantic context that enhances the distinction between instruments and tissues. The RL module plays a pivotal role in dynamically refining predictions through iterative action-space adjustments. We evaluated CLIP-RL on the EndoVis 2018 and EndoVis 2017 datasets. CLIP-RL achieved a mean IoU of 81%, outperforming state-of-the-art models, and a mean IoU of 74.12% on EndoVis 2017. This superior performance was achieved due to the combination of contrastive learning with reinforcement learning and curriculum learning.

手术分割对比学习强化学习医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。