用声音定位技术提升手术场景的三维动态理解。
Sound Source Localization for Spatial Mapping of Surgical Actions in Dynamic Scenes
- 通过麦克风阵列与深度相机融合,将声音定位投影到动态点云上。
- 在真实手术场景中实现高精度3D声源定位与多模态数据融合。
- 适合智能手术系统开发人员和医学影像研究者参考。
目的:手术场景理解是推动计算机辅助和智能手术系统发展的关键。现有方法主要依赖视觉数据或端到端学习,限制了细粒度上下文建模。本文旨在通过整合3D声学信息,增强手术场景表征,实现对动态手术环境的时序与空间感知的多模态理解。方法:提出一种新框架,通过将相控阵麦克风阵列的声源定位信息投影到RGB-D相机生成的动态点云上,构建4D音视频表示。基于Transformer的声学事件检测模块识别包含器械-组织交互的时间段,并在音视频场景表示中进行空间定位。系统在专家模拟手术过程中于真实手术室环境中进行实验评估。结果:所提方法成功实现手术声学事件在3D空间中的定位,并与视觉元素关联。实验评估表明,该方法具备准确的空间声音定位能力及鲁棒的多模态数据融合,提供了全面、动态的手术活动表征。结论:本工作首次提出动态手术场景中的空间声源定位方法,标志着多模态手术场景表征的重要进展。通过融合声学与视觉数据,该框架实现了更丰富的上下文理解,为未来智能与自主手术系统奠定基础。
原文摘要 · Abstract (English)
Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual modeling. This work aims to enhance surgical scene representations by integrating 3D acoustic information, enabling temporally and spatially aware multimodal understanding of surgical environments. Methods: We propose a novel framework for generating 4D audio-visual representations of surgical scenes by projecting acoustic localization information from a phased microphone array onto dynamic point clouds from an RGB-D camera. A transformer-based acoustic event detection module identifies relevant temporal segments containing tool-tissue interactions which are spatially localized in the audio-visual scene representation. The system was experimentally evaluated in a realistic operating room setup during simulated surgical procedures performed by experts. Results: The proposed method successfully localizes surgical acoustic events in 3D space and associates them with visual scene elements. Experimental evaluation demonstrates accurate spatial sound localization and robust fusion of multimodal data, providing a comprehensive, dynamic representation of surgical activity. Conclusion: This work introduces the first approach for spatial sound localization in dynamic surgical scenes, marking a significant advancement toward multimodal surgical scene representations. By integrating acoustic and visual data, the proposed framework enables richer contextual understanding and provides a foundation for future intelligent and autonomous surgical systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。