用大模型提升支气管镜定位的泛化与鲁棒性,实测成功率显著提高。
Harnessing Foundation Models for Robust and Generalizable 6-DOF Bronchoscopy Localization
- 融合深度估计、关键点检测与中心线约束,统一优化姿态概率。
- 在10名患者数据上,轨迹误差小于5mm的占比提升18.1%。
- 适合临床部署,对遮挡和运动模糊有强适应能力。
基于视觉的6-DOF支气管镜定位为精准、低成本的介入引导提供了可行方案。然而,现有方法存在两大瓶颈:一是在患者间泛化能力差,因标注数据稀缺;二是在视觉退化下表现不佳,因支气管镜过程常出现遮挡、运动模糊等干扰。为此,我们提出PANSv2,一种具备泛化性与鲁棒性的支气管镜定位框架。受PANS启发,该框架整合深度估计、关键点检测与中心线约束,构建统一的姿态优化机制,评估姿态概率并求解最优支气管镜位姿。为增强泛化能力,采用端镜基础模型EndoOmni进行深度估计,视频基础模型EndoMamba进行关键点检测,结合空间与时间分析。两者均在多样化内镜数据集上预训练,提供稳定可迁移的视觉表征。此外,引入自动重初始化模块,在检测到跟踪失败后,利用清晰视图下的关键点重建位姿。在包含10名患者病例的支气管镜数据集上,PANSv2实现最高跟踪成功率,相较现有方法在SR-5(绝对轨迹误差小于5mm的比例)上提升18.1%,展现出向真实临床应用转化的潜力。
原文摘要 · Abstract (English)
Vision-based 6-DOF bronchoscopy localization offers a promising solution for accurate and cost-effective interventional guidance. However, existing methods struggle with 1) limited generalization across patient cases due to scarce labeled data, and 2) poor robustness under visual degradation, as bronchoscopy procedures frequently involve artifacts such as occlusions and motion blur that impair visual information. To address these challenges, we propose PANSv2, a generalizable and robust bronchoscopy localization framework. Motivated by PANS that leverages multiple visual cues for pose likelihood measurement, PANSv2 integrates depth estimation, landmark detection, and centerline constraints into a unified pose optimization framework that evaluates pose probability and solves for the optimal bronchoscope pose. To further enhance generalization capabilities, we leverage the endoscopic foundation model EndoOmni for depth estimation and the video foundation model EndoMamba for landmark detection, incorporating both spatial and temporal analyses. Pretrained on diverse endoscopic datasets, these models provide stable and transferable visual representations, enabling reliable performance across varied bronchoscopy scenarios. Additionally, to improve robustness to visual degradation, we introduce an automatic re-initialization module that detects tracking failures and re-establishes pose using landmark detections once clear views are available. Experimental results on bronchoscopy dataset encompassing 10 patient cases show that PANSv2 achieves the highest tracking success rate, with an 18.1% improvement in SR-5 (percentage of absolute trajectory error under 5 mm) compared to existing methods, showing potential towards real clinical usage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。