医生说话即可调用视频手术导航,无需额外设备。
Speak, Segment, Track, Navigate: An Interactive System for Video-Guided Skull-Base Surgery
- 通过语音指令实时分析手术视频,自动分割并追踪器械位置。
- 工具尖端定位误差仅2.32毫米,可替代传统光学追踪系统。
- 两分钟内完成初始化,适合快速部署的微创手术场景。
我们提出一种语音引导的具身智能体框架,用于视频引导的颅底手术,能够根据外科医生的提问动态执行感知与影像导航任务。该系统将自然语言交互与实时视觉感知直接集成于术中视频流,使医生可在不中断操作的情况下请求计算辅助。与依赖外部光学追踪仪和额外硬件的传统导航系统不同,本框架仅基于术中视频运行。系统首先交互式分割并标注手术器械,将其作为空间锚点在视频流中自主追踪,支撑后续流程:解剖结构分割、术前三维模型的交互式配准、单目视频估计器械位姿,以及通过实时解剖叠加实现影像引导。我们在视频引导颅底手术场景中评估该系统,并与商用光学追踪系统对比跟踪性能。在三次实验中,混合视觉方法在相机坐标系下的工具尖端平均绝对位置误差为2.32 ± 1.10毫米,帧间偏航和俯仰传播偏差分别为0.18 ± 0.25°和0.21 ± 0.30°。系统在约两分钟内完成工具分割与解剖配准,显著降低传统追踪流程的设置复杂度。结果表明,语音引导的具身智能体可在保证高精度空间引导的同时,提升工作流整合性,实现视频引导手术系统的快速部署。
原文摘要 · Abstract (English)
We introduce a speech-guided embodied agent framework for video-guided skull base surgery that dynamically executes perception and image-guidance tasks in response to surgeon queries. The proposed system integrates natural language interaction with real-time visual perception directly on live intraoperative video streams, thereby enabling surgeons to request computational assistance without disengaging from operative tasks. Unlike conventional image-guided navigation systems that rely on external optical trackers and additional hardware setup, the framework operates purely on intraoperative video. The system begins with interactive segmentation and labeling of the surgical instrument. The segmented instrument is then used as a spatial anchor that is autonomously tracked in the video stream to support downstream workflows, including anatomical segmentation, interactive registration of preoperative 3D models, monocular video-based estimation of the surgical tool pose, and image guidance through real-time anatomical overlays. We evaluate the proposed system in video-guided skull base surgery scenarios and benchmark its tracking performance against a commercially available optical tracking system. Across three experimental trials, the hybrid vision-based method achieved a mean absolute tool-tip position error of 2.32 Plus/Minus 1.10 mm in the camera frame, with inter-frame yaw and pitch propagation discrepancies of 0.18 Plus/Minus 0.25° and 0.21 Plus/Minus 0.30°, respectively. The system completes tool segmentation and anatomy registration within approximately two minutes, substantially reducing setup complexity relative to conventional tracking workflows. These results demonstrate that speech-guided embodied agents can provide accurate spatial guidance while improving workflow integration and enabling rapid deployment of video-guided surgical systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。