arXiv:2503.02579cs.CV2025-03CVPR被引 38

构建首个多模态手术室数据集,支持全景理解与场景图生成。

MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments

  • 采集RGB-D、音频、语音、机器人日志等多模态数据
  • 包含全景分割与语义场景图标注,覆盖复杂手术场景
  • 适配医疗辅助与高危环境分析研究者使用

手术室是复杂且高风险的环境,需精准理解医护人员、器械与设备间的交互以提升手术辅助、情境感知与患者安全。现有数据集在规模、真实性和多模态表征方面存在不足,限制了手术室建模进展。为此,我们提出MM-OR,首个真实、大规模的多模态时空手术室数据集,也是首个支持多模态场景图生成的数据集。该数据集涵盖RGB-D影像、细节视图、音频、语音转录、机器人日志及追踪数据,并配有全景分割、语义场景图及下游任务标签。我们进一步提出MM2SG——首个用于场景图生成的多模态大视觉语言模型,实验证明其能有效融合多模态输入。MM-OR与MM2SG共同建立手术室整体理解的新基准,为复杂高危环境中多模态分析开辟路径。代码与数据已开源。

原文摘要 · Abstract (English)

Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety. Current datasets fall short in scale, realism and do not capture the multimodal nature of OR scenes, limiting progress in OR modeling. To this end, we introduce MM-OR, a realistic and large-scale multimodal spatiotemporal OR dataset, and the first dataset to enable multimodal scene graph generation. MM-OR captures comprehensive OR scenes containing RGB-D data, detail views, audio, speech transcripts, robotic logs, and tracking data and is annotated with panoptic segmentations, semantic scene graphs, and downstream task labels. Further, we propose MM2SG, the first multimodal large vision-language model for scene graph generation, and through extensive experiments, demonstrate its ability to effectively leverage multimodal inputs. Together, MM-OR and MM2SG establish a new benchmark for holistic OR understanding, and open the path towards multimodal scene analysis in complex, high-stakes environments. Our code, and data is available at https://github.com/egeozsoy/MM-OR.

多模态手术室场景图数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。