提升内镜图像的几何一致性,增强导航精度
Geometry-Consistent Endoscopic Representations for Image-Guided Navigation via Structured Foundation Model Adaptation

- 用合成数据+分层适配器,让模型学出更准的几何特征
- 在多个内镜任务中,姿态估计和深度预测性能显著提升
- 适合做内镜导航的模型初始化,尤其适用于数据少的场景
单目内镜视觉导航因深度线索有限、组织纹理弱、非刚性形变及跨域外观差异大而困难,导致位姿估计、深度预测与图像-解剖对齐复杂。尽管近期视觉基础模型表现良好,其表征仍缺乏足够的几何一致性,影响特征对应稳定性,限制下游导航可靠性。本文提出统一框架,学习几何一致且领域鲁棒的单目内镜表征。该框架结合提供精确几何监督的合成数据管道,以及分层感知的几何-语义适配机制——一种结构化替代标准LoRA的方法,在Transformer层级中选择性插入低秩适配器,并耦合逐层训练目标,以促进中间特征的几何对应与深层特征的语义一致性。在公开与私有数据集上的实验表明,所学表征几何与语义质量更高,下游导航任务(如位姿估计、单目深度估计)性能更好。表征在临床支气管镜中表现出优异的合成到真实迁移能力,并为有限监督下的鼻窦镜和结肠镜适应提供有效初始化。框架还展现出良好的模型规模与数据量扩展性。结果支持分层感知、几何引导的适配是内镜表征学习的可行方案。
原文摘要 · Abstract (English)
Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation, and substantial appearance variation across domains, all of which complicate pose estimation, depth prediction, and image-to-anatomy alignment. Although recent vision foundation models have shown promise, their learned representations often remain insufficiently geometry-consistent, hindering stable feature correspondence and limiting their reliability for downstream navigation tasks. We propose a unified framework for learning geometry-consistent and domain-robust image representations for monocular endoscopy. The framework combines a synthetic data pipeline that provides accurate geometric supervision with Hierarchy-Aware Geometry-Semantic Adaptation, a structured alternative to standard LoRA that inserts low-rank adapters selectively across the transformer hierarchy and couples them with layer-wise training objectives to encourage geometric correspondence in intermediate features and semantic consistency in deeper features. Experiments on public and proprietary datasets show improved geometric and semantic representation quality, leading to better performance on downstream navigation tasks including pose estimation and monocular depth estimation. The learned representations show favorable synthetic-to-real transfer on clinical bronchoscopy and provide a useful initialization for adaptation to sinus endoscopy and colonoscopy under limited supervision. The framework also shows favorable scaling with model size and training data. These results support hierarchy-aware, geometry-guided adaptation as a practical approach for endoscopic representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。