arXiv:2503.21054eess.IVcs.CV2025-03被引 16

用数字孪生技术让手术室流程分析更灵活精准

Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins

  • 构建手术室数字孪生模型,保留物体语义与空间关系
  • 无需微调大模型,实现跨场景手术流程推理分割,准确率提升6.12%-9.74%
  • 支持自然语言查询,自动生成带视觉证据的分析报告,适合医疗管理者和科研人员

分析手术室(OR)工作流程以获得关于效率的定量洞察,对医院提升患者护理水平和财务可持续性至关重要。以往的OR级工作流程分析依赖端到端深度神经网络,虽在受限环境中表现良好,但仅适用于开发时设定的条件,难以适应不同场景(如大型学术中心与乡村医疗机构)的需求,且需重新采集数据、标注与训练。基于基础模型的推理分割(RS)可提供灵活性,仅通过与目标对象相关的隐式文本查询即可自动分析来自手术室视频流的工作流程。然而,现有RS方法依赖大语言模型(LLM)微调,在理解语义/空间关系方面表现不佳,且因视觉特征差异和领域术语变化,泛化能力有限。为此,我们首先提出一种新型数字孪生(DT)表示,保留各手术室组件间的语义与空间关系;在此基础上,提出无需微调LLM的ORS-DS框架,将推理分割重构为“推理-检索-合成”范式;最后,设计基于LLM的ORDiRS-Agent,将复杂查询分解为可管理的子查询,并结合详细文字解释与来自推理分割的视觉证据生成响应。在自建及公开的两个手术室数据集上的实验表明,我们的ORDiRS相比现有最先进方法在cIoU上提升6.12%-9.74%。

原文摘要 · Abstract (English)

Analyzing operating room (OR) workflows to derive quantitative insights into OR efficiency is important for hospitals to maximize patient care and financial sustainability. Prior work on OR-level workflow analysis has relied on end-to-end deep neural networks. While these approaches work well in constrained settings, they are limited to the conditions specified at development time and do not offer the flexibility necessary to accommodate the OR workflow analysis needs of various OR scenarios (e.g., large academic center vs. rural provider) without data collection, annotation, and retraining. Reasoning segmentation (RS) based on foundation models offers this flexibility by enabling automated analysis of OR workflows from OR video feeds given only an implicit text query related to the objects of interest. Due to the reliance on large language model (LLM) fine-tuning, current RS approaches struggle with reasoning about semantic/spatial relationships and show limited generalization to OR video due to variations in visual characteristics and domain-specific terminology. To address these limitations, we first propose a novel digital twin (DT) representation that preserves both semantic and spatial relationships between the various OR components. Then, building on this foundation, we propose ORDiRS (Operating Room Digital twin representation for Reasoning Segmentation), an LLM-tuning-free RS framework that reformulates RS into a "reason-retrieval-synthesize" paradigm. Finally, we present ORDiRS-Agent, an LLM-based agent that decomposes OR workflow analysis queries into manageable RS sub-queries and generates responses by combining detailed textual explanations with supporting visual evidence from RS. Experimental results on both an in-house and a public OR dataset demonstrate that our ORDiRS achieves a cIoU improvement of 6.12%-9.74% compared to the existing state-of-the-arts.

数字孪生手术室分析推理分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。