arXiv:2603.01108cs.CV2026-03

首个语言引导的手术器械定位基准,支持精准识别具体器械。

GroundedSurg: A Multi-Procedure Benchmark for Language-Conditioned Surgical Tool Segmentation

  • 用自然语言描述指定手术器械,结合空间标注进行实例级定位。
  • 涵盖四大类手术,包含多种器械与复杂场景,覆盖真实临床需求。
  • 适合研究视觉语言模型在手术辅助中的定位与理解能力。

临床可靠的手术场景感知对智能、情境感知的术中辅助(如器械交接指引、碰撞规避、流程自适应机器人支持)至关重要。现有手术器械基准主要评估类别级分割,要求模型检测预定义器械类别的所有实例。然而,真实临床决策常需基于器械的功能角色、空间关系或解剖交互能力来定位特定实例,当前评估范式无法捕捉此类需求。我们提出 GroundedSurg,首个语言条件下的实例级手术定位基准。每个样本包含一张手术图像与针对单一器械的自然语言描述,并附带结构化空间标注(边界框与点级锚点)。数据集涵盖眼科、腹腔镜、机器人及开放手术,覆盖多样器械类型、成像条件与操作复杂度。通过联合评估语言引用解析与像素级定位,GroundedSurg实现对视觉-语言模型在多器械临床场景中表现的系统性、真实性评估。大量实验显示现代分割模型与视觉语言模型间存在显著性能差距,凸显手术AI系统中临床语义感知推理的紧迫性。代码与数据公开于 https://github.com/gaash-lab/GroundedSurg。

原文摘要 · Abstract (English)

Clinically reliable perception of surgical scenes is essential for advancing intelligent, context-aware intraoperative assistance such as instrument handoff guidance, collision avoidance, and workflow-aware robotic support. Existing surgical tool benchmarks primarily evaluate category-level segmentation, requiring models to detect all instances of predefined instrument classes. However, real-world clinical decisions often require resolving references to a specific instrument instance based on its functional role, spatial relation, or anatomical interaction capabilities not captured by current evaluation paradigms. We introduce GroundedSurg, the first language-conditioned, instance-level surgical grounding benchmark. Each instance pairs a surgical image with a natural-language description targeting a single instrument, accompanied by structured spatial grounding annotations including bounding boxes and point-level anchors. The dataset spans ophthalmic, laparoscopic, robotic, and open procedures, encompassing diverse instrument types, imaging conditions, and operative complexities. By jointly evaluating linguistic reference resolution and pixel-level localization, GroundedSurg enables a systematic and realistic evaluation of vision-language models in clinically realistic multi-instrument scenes. Extensive experiments demonstrate substantial performance gaps across modern segmentation and VLMs, highlighting the urgent need for clinically grounded vision-language reasoning in surgical AI systems. Code and data are publicly available at https://github.com/gaash-lab/GroundedSurg

手术分割视觉语言医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。