arXiv:2602.09252cs.CVcs.AI2026-02

用自然语言指导手术图像分割,支持医生互动并持续优化结果。

VLM-Guided Iterative Refinement for Surgical Image Segmentation with Foundation Models

  • 通过视觉语言模型和自适应流程,实现分割结果的迭代优化。
  • 在内域与跨域数据上均达顶尖性能,医生反馈进一步提升效果。
  • 首个支持自然语言交互的手术分割自修正框架,适合临床协同场景。

手术图像分割对机器人辅助手术和术中导航至关重要。现有方法受限于预定义类别,仅生成一次性预测,缺乏自适应优化机制,且无法支持临床医生交互。本文提出IR-SIS系统,利用微调后的SAM3进行初始分割,借助视觉语言模型检测器械并评估分割质量,采用代理式工作流自适应选择优化策略。系统支持通过自然语言反馈的医生在环交互。同时构建了基于EndoVis2017和EndoVis2018基准的多粒度语言标注数据集。实验表明,在同分布与跨分布数据上均达到当前最优性能,医生参与带来额外提升。本工作首次建立具备自适应迭代优化能力的语言驱动手术分割框架。

原文摘要 · Abstract (English)

Surgical image segmentation is essential for robot-assisted surgery and intraoperative guidance. However, existing methods are constrained to predefined categories, produce one-shot predictions without adaptive refinement, and lack mechanisms for clinician interaction. We propose IR-SIS, an iterative refinement system for surgical image segmentation that accepts natural language descriptions. IR-SIS leverages a fine-tuned SAM3 for initial segmentation, employs a Vision-Language Model to detect instruments and assess segmentation quality, and applies an agentic workflow that adaptively selects refinement strategies. The system supports clinician-in-the-loop interaction through natural language feedback. We also construct a multi-granularity language-annotated dataset from EndoVis2017 and EndoVis2018 benchmarks. Experiments demonstrate state-of-the-art performance on both in-domain and out-of-distribution data, with clinician interaction providing additional improvements. Our work establishes the first language-based surgical segmentation framework with adaptive self-refinement capabilities.

医学图像视觉语言自适应分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。