arXiv:2601.12224cs.CVcs.AI2026-01AAAI被引 2

用工具运动轨迹实现语言驱动的手术器械分割,提升复杂场景下的识别能力。

Where It Moves, It Matters: Referring Surgical Instrument Segmentation via Motion

  • 基于器械运动轨迹而非外观进行语言引导分割
  • 在遮挡和陌生术语下仍保持高准确率,平均mIoU达68.7%
  • 适合开发智能手术助手与自主机器人系统

实现自然语言驱动的手术场景交互是智能手术室和自主手术机器人辅助的关键一步。然而,基于自然语言描述定位手术器械的指代分割任务在手术视频中仍研究不足,现有方法因依赖静态视觉线索和预定义器械名称而泛化能力差。本文提出SurgRef,一种新型运动引导框架,通过捕捉工具随时间的运动与交互行为,而非其外观特征,实现自由形式语言表达的精准定位。该方法可在遮挡、模糊或使用非标准术语时仍有效识别器械。为训练与评估SurgRef,我们构建了Ref-IMotion数据集,涵盖多机构、多样化的手术视频,提供密集时空标注和以运动为中心的丰富语言表达。SurgRef在多种手术流程中均达到当前最优性能,建立了鲁棒、语言驱动手术视频分割的新基准。

原文摘要 · Abstract (English)

Enabling intuitive, language-driven interaction with surgical scenes is a critical step toward intelligent operating rooms and autonomous surgical robotic assistance. However, the task of referring segmentation, localizing surgical instruments based on natural language descriptions, remains underexplored in surgical videos, with existing approaches struggling to generalize due to reliance on static visual cues and predefined instrument names. In this work, we introduce SurgRef, a novel motion-guided framework that grounds free-form language expressions in instrument motion, capturing how tools move and interact across time, rather than what they look like. This allows models to understand and segment instruments even under occlusion, ambiguity, or unfamiliar terminology. To train and evaluate SurgRef, we present Ref-IMotion, a diverse, multi-institutional video dataset with dense spatiotemporal masks and rich motion-centric expressions. SurgRef achieves state-of-the-art accuracy and generalization across surgical procedures, setting a new benchmark for robust, language-driven surgical video segmentation.

手术分割语言引导运动建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。