融合显微镜与OCT图像,实现实时精准器械追踪
Towards Comprehensive Real-Time Scene Understanding in Ophthalmic Surgery through Multimodal Image Fusion
- 用交叉注意力融合显微镜与OCT特征,支持多任务联合预测
- 器械定位准确率95.79% mAP50,视网膜附近距离误差降至33μm
- 适合需要高精度实时导航的玻璃体视网膜手术场景
目的:将多模态成像融入手术室,可实现全面的术中场景理解。在眼科手术中,目前有两种互补的成像方式:术中显微镜(OPMI)和实时术中光学相干断层扫描(iOCT)。本文首次探索了时间维度上OPMI与iOCT特征融合,展示了多模态图像处理在多头预测中的潜力,以玻璃体视网膜手术中的精确器械追踪为例。方法:提出一种多模态、时序、支持实时处理的网络架构,实现联合器械检测、关键点定位与工具-组织距离估计。网络通过YoloNAS和CNN编码器分别高效提取OPMI与iOCT特征,并引入交叉注意力融合模块进行特征融合;此外,基于区域的递归模块利用时间一致性。结果:实验表明,器械定位与关键点检测表现可靠(mAP50达95.79%),且引入iOCT显著提升工具-组织距离估计性能,达到每帧22.5毫秒的实时处理速度。尤其在距视网膜小于1毫米的近距离下,距离估计误差由仅使用OPMI时的284μm降低至33μm。结论:多模态特征融合相比单模态处理能提升多任务预测精度,通过定制化网络设计可实现实时性能。尽管本研究展示了多模态处理在影像引导玻璃体视网膜手术中的潜力,也凸显了未来研究需解决的关键挑战,以实现更可靠、一致和全面的术中场景理解。
原文摘要 · Abstract (English)
Purpose: The integration of multimodal imaging into operating rooms paves the way for comprehensive surgical scene understanding. In ophthalmic surgery, by now, two complementary imaging modalities are available: operating microscope (OPMI) imaging and real-time intraoperative optical coherence tomography (iOCT). This first work toward temporal OPMI and iOCT feature fusion demonstrates the potential of multimodal image processing for multi-head prediction through the example of precise instrument tracking in vitreoretinal surgery. Methods: We propose a multimodal, temporal, real-time capable network architecture to perform joint instrument detection, keypoint localization, and tool-tissue distance estimation. Our network design integrates a cross-attention fusion module to merge OPMI and iOCT image features, which are efficiently extracted via a YoloNAS and a CNN encoder, respectively. Furthermore, a region-based recurrent module leverages temporal coherence. Results: Our experiments demonstrate reliable instrument localization and keypoint detection (95.79% mAP50) and show that the incorporation of iOCT significantly improves tool-tissue distance estimation, while achieving real-time processing rates of 22.5 ms per frame. Especially for close distances to the retina (below 1 mm), the distance estimation accuracy improved from 284 $μm$ (OPMI only) to 33 $μm$ (multimodal). Conclusion: Feature fusion of multimodal imaging can enhance multi-task prediction accuracy compared to single-modality processing and real-time processing performance can be achieved through tailored network design. While our results demonstrate the potential of multi-modal processing for image-guided vitreoretinal surgery, they also underline key challenges that motivate future research toward more reliable, consistent, and comprehensive surgical scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。