arXiv:2509.10748cs.CV2025-09被引 1

用语音协同实现术中场景的实时分割与追踪,无需预设标签。

SCOPE: Speech-guided COllaborative PErception Framework for Surgical Scene Segmentation

  • 结合大模型推理与开放集视觉模型,通过语音反馈动态生成分割结果。
  • 在Cataract1k和自建颅底数据集上实现术中器械与解剖结构的实时分割与追踪。
  • 支持医生自然语音交互,适合需要灵活适应新场景的智能手术辅助系统。

准确分割与追踪手术场景中的相关元素对于实现术中上下文感知辅助与决策至关重要。现有方法仍依赖特定领域、有监督模型,需标注数据且难以扩展至新场景或预定义类别之外。近期基于提示的视觉基础模型(VFM)实现了跨异构医学图像的开放集零样本分割。然而,这些模型对人工视觉或文本提示的依赖限制了其在术中环境的部署。本文提出一种语音引导的协作感知框架(SCOPE),融合大语言模型(LLM)的推理能力与开放集视觉基础模型(VFM)的感知能力,支持在术中视频流中实时进行器械与解剖结构的分割、标注与追踪。框架核心是一个协作感知代理,能生成VFM输出的候选分割结果,并结合临床医生的直观语音反馈,以自然人机协作方式引导器械分割。随后,器械自身作为交互指针,用于标注手术场景中的其他元素。我们在公开的Cataract1k数据集子集及自建的体外颅底数据集上评估该框架,验证其在术中实现动态分割与追踪的潜力。此外,通过一次现场模拟体外实验展示了其动态适应能力。该人机协作范式为开发可适配、免手动、以外科医生为中心的动态手术室工具提供了可能。

原文摘要 · Abstract (English)

Accurate segmentation and tracking of relevant elements of the surgical scene is crucial to enable context-aware intraoperative assistance and decision making. Current solutions remain tethered to domain-specific, supervised models that rely on labeled data and required domain-specific data to adapt to new surgical scenarios and beyond predefined label categories. Recent advances in prompt-driven vision foundation models (VFM) have enabled open-set, zero-shot segmentation across heterogeneous medical images. However, dependence of these models on manual visual or textual cues restricts their deployment in introperative surgical settings. We introduce a speech-guided collaborative perception (SCOPE) framework that integrates reasoning capabilities of large language model (LLM) with perception capabilities of open-set VFMs to support on-the-fly segmentation, labeling and tracking of surgical instruments and anatomy in intraoperative video streams. A key component of this framework is a collaborative perception agent, which generates top candidates of VFM-generated segmentation and incorporates intuitive speech feedback from clinicians to guide the segmentation of surgical instruments in a natural human-machine collaboration paradigm. Afterwards, instruments themselves serve as interactive pointers to label additional elements of the surgical scene. We evaluated our proposed framework on a subset of publicly available Cataract1k dataset and an in-house ex-vivo skull-base dataset to demonstrate its potential to generate on-the-fly segmentation and tracking of surgical scene. Furthermore, we demonstrate its dynamic capabilities through a live mock ex-vivo experiment. This human-AI collaboration paradigm showcase the potential of developing adaptable, hands-free, surgeon-centric tools for dynamic operating-room environments.

手术分割语音交互协作感知零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。