用合成数据训练模型,实现精准肺镜导航定位。
BronchOpt : Vision-Based Pose Optimization with Fine-Tuned Foundation Models for Accurate Bronchoscopy Navigation
- 通过微调跨模态编码器直接比对内镜图像与CT深度图。
- 在真实患者数据上达2.65毫米平移误差,0.19弧度旋转误差。
- 首个公开合成数据集,推动肺镜导航可复现评估。
术中将支气管镜尖端精确定位到患者解剖结构仍面临呼吸运动、解剖差异及CT与实际人体间的形变错位等挑战。现有视觉方法泛化能力差,导致残余对齐误差。本文提出一种基于视觉的姿态优化框架,实现术中内镜图像与术前CT之间的逐帧2D-3D配准。通过微调的模态与领域不变编码器,直接计算真实内镜RGB图像与CT渲染深度图间的相似性;结合可微渲染模块,通过深度一致性迭代优化相机姿态。为提升可复现性,引入首个公开的支气管镜导航合成基准数据集,解决缺乏配对CT-内镜数据的问题。模型仅在合成数据上训练,却在基准上实现平均平移误差2.65毫米、旋转误差0.19弧度,表现稳定准确。真实患者数据上的定性结果表明其具备强跨域泛化能力,无需领域特定适配即可实现一致的逐帧2D-3D对齐。整体框架通过迭代视觉优化实现鲁棒、领域无关的定位,新基准则为视觉支气管镜导航的标准化进展奠定基础。
原文摘要 · Abstract (English)
Accurate intra-operative localization of the bronchoscope tip relative to patient anatomy remains challenging due to respiratory motion, anatomical variability, and CT-to-body divergence that cause deformation and misalignment between intra-operative views and pre-operative CT. Existing vision-based methods often fail to generalize across domains and patients, leading to residual alignment errors. This work establishes a generalizable foundation for bronchoscopy navigation through a robust vision-based framework and a new synthetic benchmark dataset that enables standardized and reproducible evaluation. We propose a vision-based pose optimization framework for frame-wise 2D-3D registration between intra-operative endoscopic views and pre-operative CT anatomy. A fine-tuned modality- and domain-invariant encoder enables direct similarity computation between real endoscopic RGB frames and CT-rendered depth maps, while a differentiable rendering module iteratively refines camera poses through depth consistency. To enhance reproducibility, we introduce the first public synthetic benchmark dataset for bronchoscopy navigation, addressing the lack of paired CT-endoscopy data. Trained exclusively on synthetic data distinct from the benchmark, our model achieves an average translational error of 2.65 mm and a rotational error of 0.19 rad, demonstrating accurate and stable localization. Qualitative results on real patient data further confirm strong cross-domain generalization, achieving consistent frame-wise 2D-3D alignment without domain-specific adaptation. Overall, the proposed framework achieves robust, domain-invariant localization through iterative vision-based optimization, while the new benchmark provides a foundation for standardized progress in vision-based bronchoscopy navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。