arXiv:2411.12199cs.CV2024-11被引 1

提出新任务R-SIS,让模型在不知道器械是否存在时仍能准确响应文本提示。

Rethinking Text-Promptable Surgical Instrument Segmentation with Robust Framework

  • 设计新框架R-SIS,允许在未知器械存在情况下通过文本提示选择性分割
  • 实测现有方法在无真实器械时仍大量误检,误报率显著
  • 适合医疗视觉系统开发、机器人手术自动化等需应对不确定性的场景

外科器械分割是计算机辅助与机器人手术系统的关键组成部分。基于视觉的分割模型通常只能输出预定义类别中的器械,限制了其在交互式系统和机器人任务自动化中的应用。可提示分割方法虽能根据文本提示进行选择性预测,但普遍假设场景中出现的器械已知,导致无法泛化到未见过或动态出现的器械。在实际手术环境中,器械存在信息往往不可用,此假设不成立,引发大量假阳性分割。为此,本文提出新任务——鲁棒文本提示外科器械分割(R-SIS):在不知晓器械是否存在的前提下,对所有候选类别发出提示,仅当器械实际可见时才生成掩码。该设定更贴近真实手术场景中的不确定性。我们在多个外科视频数据集上评估现有方法在R-SIS协议下的表现,发现缺少真实器械时仍存在显著假阳性预测。结果表明当前评估协议与真实应用存在脱节,亟需考虑提示不确定性和器械缺失的新基准。

原文摘要 · Abstract (English)

Surgical instrument segmentation is an essential component of computer-assisted and robotic surgery systems. Vision-based segmentation models typically produce outputs limited to a predefined set of instrument categories, which restricts their applicability in interactive systems and robotic task automation. Promptable segmentation methods allow selective predictions based on textual prompts. However, they often rely on the assumption that the instruments present in the scene are already known, and prompts are generated accordingly, limiting their ability to generalize to unseen or dynamically emerging instruments. In practical surgical environments, where instrument existence information is not provided, this assumption does not hold consistently, resulting in false-positive segmentation. To address these limitations, we formulate a new task called Robust text-promptable Surgical Instrument Segmentation (R-SIS). Under this setting, prompts are issued for all candidate categories without access to instrument presence information. R-SIS requires distinguishing which prompts refer to visible instruments and generating masks only when such instruments are explicitly present in the scene. This setting reflects practical conditions where uncertainty in instrument presence is inherent. We evaluate existing segmentation methods under the R-SIS protocol using surgical video datasets and observe substantial false-positive predictions in the absence of ground-truth instruments. These findings demonstrate a mismatch between current evaluation protocols and real-world use cases, and support the need for benchmarks that explicitly account for prompt uncertainty and instrument absence.

医学图像分割提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。