首个融合第一/第三人称视角的手术室数据集,助力智能手术感知。
EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
- 结合穿戴设备与多视角摄像头,融合第一人称与第三人称数据。
- 包含94分钟、84,553帧的手术视频,标注56万+三元组场景关系。
- 适合研究手术行为理解、多模态感知的科研人员使用。
手术室环境快速且遮挡严重,需要先进感知模型提升安全与效率。现有数据集或仅有部分第一人称视角,或缺少丰富的第三人称多视角信息,未充分融合双重视角。我们提出EgoExOR,首个融合第一人称与第三人称视角的手术室数据集及基准评测体系。该数据集涵盖94分钟(84,553帧,15 FPS)的两例模拟脊柱手术——超声引导穿刺与微创脊柱手术,集成可穿戴眼镜获取的RGB、注视点、手部追踪与音频数据,以及多视角RGB-D相机采集的RGB与深度图,还有超声影像。其详尽的场景图标注覆盖36个实体与22种关系(共568,235个三元组),支持动作识别与以人为核心的感知任务建模。我们评估了两种改进的前沿模型在手术场景图生成上的表现,并提供一个显式利用多模态、多视角信号的新基线。该数据集与基准为手术室感知研究奠定新基础,提供丰富多模态资源,推动下一代临床感知技术发展。
原文摘要 · Abstract (English)
Operating rooms (ORs) demand precise coordination among surgeons, nurses, and equipment in a fast-paced, occlusion-heavy environment, necessitating advanced perception models to enhance safety and efficiency. Existing datasets either provide partial egocentric views or sparse exocentric multi-view context, but do not explore the comprehensive combination of both. We introduce EgoExOR, the first OR dataset and accompanying benchmark to fuse first-person and third-person perspectives. Spanning 94 minutes (84,553 frames at 15 FPS) of two emulated spine procedures, Ultrasound-Guided Needle Insertion and Minimally Invasive Spine Surgery, EgoExOR integrates egocentric data (RGB, gaze, hand tracking, audio) from wearable glasses, exocentric RGB and depth from RGB-D cameras, and ultrasound imagery. Its detailed scene graph annotations, covering 36 entities and 22 relations (568,235 triplets), enable robust modeling of clinical interactions, supporting tasks like action recognition and human-centric perception. We evaluate the surgical scene graph generation performance of two adapted state-of-the-art models and offer a new baseline that explicitly leverages EgoExOR's multimodal and multi-perspective signals. This new dataset and benchmark set a new foundation for OR perception, offering a rich, multimodal resource for next-generation clinical perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。