arXiv:2505.18381cs.CV2025-05

无需标记点和追踪设备,用单目镜头实现耳蜗手术实时定位

Monocular Marker-free Patient-to-Image Intraoperative Registration for Cochlear Implant Surgery

  • 基于合成数据训练轻量网络,直接映射术前CT到术中图像
  • 9例临床验证,多数情况姿态角误差小于10度,满足临床精度
  • 兼容单目显微镜,无额外硬件依赖,适合实际手术场景

本文提出一种新型单目术中患者-图像配准方法,专为耳蜗植入手术设计,无需外部追踪设备或标记点。利用包含广泛变换的合成显微手术数据集,通过零样本学习方法,采用轻量神经网络将术前CT扫描直接映射至术中2D图像,实现实时手术导航。该框架通过学习合成数据估计相机位姿(含旋转矩阵与平移向量),实现精准高效配准。在9个临床病例中,采用个体特异性和跨患者验证策略评估,结果表明该方法在多数情况下6自由度相机位姿预测的角误差低于10度,解决了传统方法依赖外部追踪系统或标记点的局限性。

原文摘要 · Abstract (English)

This paper presents a novel method for monocular patient-to-image intraoperative registration, specifically designed to operate without any external hardware tracking equipment or fiducial point markers. Leveraging a synthetic microscopy surgical scene dataset with a wide range of transformations, our approach directly maps preoperative CT scans to 2D intraoperative surgical frames through a lightweight neural network for real-time cochlear implant surgery guidance via a zero-shot learning approach. Unlike traditional methods, our framework seamlessly integrates with monocular surgical microscopes, making it highly practical for clinical use without additional hardware dependencies and requirements. Our method estimates camera poses, which include a rotation matrix and a translation vector, by learning from the synthetic dataset, enabling accurate and efficient intraoperative registration. The proposed framework was evaluated on nine clinical cases using a patient-specific and cross-patient validation strategy. Our results suggest that our approach achieves clinically relevant accuracy in predicting 6D camera poses for registering 3D preoperative CT scans to 2D surgical scenes with an angular error within 10 degrees in most cases, while also addressing limitations of traditional methods, such as reliance on external tracking systems or fiducial markers.

手术导航单目视觉耳蜗植入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。