用视觉语言模型提升支气管镜6自由度定位精度,兼顾准确与效率
BREATH-VL: Vision-Language-Guided 6-DoF Bronchoscopy Localization via Semantic-Geometric Fusion
- 融合视觉语言模型语义与几何配准的双路定位方法
- 在真实气道数据集上实现25.5%的平移误差降低
- 轻量时序机制支持高效运动历史建模,适合临床部署
视觉语言模型(VLM)在导航与定位任务中表现优异,得益于其大规模预训练带来的语义理解能力。然而,将其应用于6-DoF内窥镜定位仍面临三大挑战:1)缺乏大规模、高质量、密集标注且面向定位的真实医疗场景视觉语言数据集;2)细粒度位姿回归能力有限;3)从前帧提取时序特征计算延迟高。为此,我们构建了当前最大的在体支气管镜定位数据集BREATH,采集自复杂的人体气道环境。基于该数据集,提出BREATH-VL混合框架,将VLM的语义线索与基于视觉的配准方法的几何信息相结合,实现精准6-DoF姿态估计。核心思想是二者优势互补:VLM提供可泛化的语义理解,配准方法实现精确几何对齐。为进一步增强VLM对时序上下文的捕捉能力,引入轻量级上下文学习机制,将运动历史编码为语言提示,实现高效时序推理而无需昂贵的视频级计算。大量实验表明,视觉语言模块在复杂手术场景中展现出鲁棒的语义定位能力。在此基础上,BREATH-VL在准确性和泛化性上均优于现有纯视觉定位方法,相比最优基线,平移误差降低25.5%,同时保持可接受的计算延迟。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have recently shown remarkable performance in navigation and localization tasks by leveraging large-scale pretraining for semantic understanding. However, applying VLMs to 6-DoF endoscopic camera localization presents several challenges: 1) the lack of large-scale, high-quality, densely annotated, and localization-oriented vision-language datasets in real-world medical settings; 2) limited capability for fine-grained pose regression; and 3) high computational latency when extracting temporal features from past frames. To address these issues, we first construct BREATH dataset, the largest in-vivo endoscopic localization dataset to date, collected in the complex human airway. Building on this dataset, we propose BREATH-VL, a hybrid framework that integrates semantic cues from VLMs with geometric information from vision-based registration methods for accurate 6-DoF pose estimation. Our motivation lies in the complementary strengths of both approaches: VLMs offer generalizable semantic understanding, while registration methods provide precise geometric alignment. To further enhance the VLM's ability to capture temporal context, we introduce a lightweight context-learning mechanism that encodes motion history as linguistic prompts, enabling efficient temporal reasoning without expensive video-level computation. Extensive experiments demonstrate that the vision-language module delivers robust semantic localization in challenging surgical scenes. Building on this, our BREATH-VL outperforms state-of-the-art vision-only localization methods in both accuracy and generalization, reducing translational error by 25.5% compared with the best-performing baseline, while achieving competitive computational latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。