用轻量定位+大模型实现超声引导下气管解剖的精准分层理解
Learning-based Hierarchical Tracheal Anatomy Understanding from Sparse Surgical Demonstration Annotations for Ultrasound Robots

- 结合YOLOv8n与稀疏提示优化的SAM2,从少量标注中实现高精度分割
- 平均Dice系数达0.777,显著优于U-Net的0.494,且保持6.92帧/秒速度
- 适合需高精度、低延迟的机器人辅助气管切开手术系统
气管切开术需要精确定位气管切口位置,但传统手动触诊主观性强且不可靠,而超声使用仍依赖操作者。本文提出一种面向超声引导机器人的分层气管解剖理解学习框架。采用两阶段感知流程,融合YOLOv8n定位主干与稀疏提示优化的SAM2解码器,实现从稀疏手术标注中高保真分割。混合训练策略连接受控实验室数据与非约束序列,确保临床鲁棒性。实验表明,该解耦架构在控制域与泛化域均保持稳定性能,平均Dice相似系数(DSC)达0.777,显著优于U-Net基线(泛化域DSC ≤ 0.494),后者常出现解剖碎片化与性能下降。通过将掩码解码限制于目标稀疏兴趣区域,模型达到6.92 FPS吞吐率,满足闭环机器人遥操作需求。研究证实,轻量级定位与基础规模视觉模型结合可实现鲁棒的分层气管解剖理解,为标准化、自主化外科辅助建立可扩展基础,有效应对真实世界超声的变异性,提升机器人辅助气管切开的安全性与精确性。
原文摘要 · Abstract (English)
Tracheostomy requires precise localization of the tracheal incision site; however, conventional manual palpation is subjective and often unreliable, while ultrasound utility remains operator-dependent. This work presents a learning-based framework for hierarchical tracheal anatomy understanding, designed specifically for ultrasound-guided robotic systems. We propose a two-stage perception pipeline integrating a YOLOv8n localization backbone with a sparse, prompt-optimized SAM2 decoder to achieve high-fidelity segmentation from sparse surgical annotations. Our hybrid training strategy, bridging curated laboratory data with unconstrained sequences, ensures clinical robustness. Experimental benchmarks demonstrate that this decoupled architecture effectively balances generalization, precision, and efficiency. The YOLOv8n and SAM2 framework achieves a consistent Mean Dice Similarity Coefficient (DSC) of 0.777 across both controlled and generalized domains. This significantly outperforms U-Net baselines, which often suffer from anatomical fragmentation and performance degradation (Generalization DSC $\le$ 0.494). By constraining mask decoding to targeted, sparse regions of interest, our model achieves a throughput of 6.92 FPS, which is vital for closed-loop robotic teleoperation. This study confirms that a robust hierarchical understanding of tracheal anatomy can be derived by coupling lightweight localization with foundation-scale visual models. Our framework establishes a scalable foundation for standardized, autonomous surgical assistance, effectively navigating the variability of real-world ultrasound to enhance the safety and precision of robotic-assisted tracheostomy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。