arXiv:2508.08199cs.CV2025-08中稿 · ACM Multimedia 202…被引 8

仅用彩色图像实现手术室3D空间理解,无需额外传感器

Generalizable Operating Room Expert with Multimodal Enhancement

  • 从单张彩色图推断深度、分割和点云信息,构建结构化空间表征
  • 在多个手术室基准上达到当前最优性能,且能泛化到未见场景
  • 适合需要低成本部署的临床空间感知系统研发人员

手术室中的精确空间建模对术中认知、风险规避和手术决策至关重要。现有方法虽利用多模态数据学习空间关系,但许多依赖难以在真实临床环境中部署的传感模态,且在传感受限条件下缺乏明确的3D推理能力。同时,主要基于易获取2D数据训练的模型往往无法捕捉复杂手术室场景的细粒度几何与语义结构。为此,我们提出 extbf{OR-Expert},一种仅使用RGB图像进行3D空间推理的大规模视觉语言框架。OR-Expert内部从RGB图像中推导出深度、全景分割和点云线索,并将其编码为结构化空间表示。其空间增强特征融合模块将这些伪模态与RGB及文本特征在共享标记空间中对齐,实现语义、几何与语言的联合推理。该统一端到端多模态大语言模型无需外部深度、分割或点云传感器,也无需推理时附加专家标注,即可支持细致的空间理解。在多个手术室基准上的实验表明,OR-Expert实现了当前最优性能,并有效泛化至未见手术场景及下游空间推理任务。

原文摘要 · Abstract (English)

Precise spatial modeling in the operating room (OR) is essential for intraoperative awareness, hazard avoidance, and surgical decision-making. Although existing approaches exploit multimodal data to learn spatial relationships, many depend on sensing modalities that are difficult to deploy in real clinical environments and remain limited in explicit 3D reasoning under constrained sensing conditions. Meanwhile, models trained primarily on readily available 2D data often fail to capture the fine-grained geometric and semantic structure of complex OR scenes. To address these limitations, we introduce \textbf{OR-Expert}, a large vision-language framework for 3D spatial reasoning with RGB-only inference. OR-Expert internally derives depth, panoptic segmentation, and point-cloud cues from RGB images and encodes them as structured spatial representations. Its Spatial-Enhanced Feature Fusion Block aligns these pseudo-modalities with RGB and textual features in a shared token space, enabling joint semantic, geometric, and language reasoning. The unified end-to-end MLLM therefore supports detailed spatial understanding without requiring external depth, segmentation, or point-cloud sensors, or additional expert annotations at inference time. Experiments on multiple operating-room benchmarks demonstrate that OR-Expert achieves state-of-the-art performance and generalizes effectively to unseen surgical scenes and downstream spatial reasoning tasks.

3D空间理解手术室智能多模态大模型视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。