arXiv:2604.27383eess.IVcs.CV2026-04

提出轻量级网络,实现实时高精度声门分割,提升鼻导管插管导航效率。

A Real-time Scale-robust Network for Glottis Segmentation in Nasal Transnasal Intubation

论文配图:A Real-time Scale-robust Network for Glottis Segmentation in Nasal Transnasal Intubation
图 1 · 摘自论文原文
  • 设计多感受野模块,增强对声门尺度变化的鲁棒性。
  • 在三个数据集上达到92.9% mDice,推理速度超170帧/秒。
  • 模型仅19MB,适合移动端实时应用,临床导航场景友好。

鼻气管插管(NTI)是维持患者气道通畅的关键临床操作。基于视觉辅助的NTI已成优化操作效率、减少人为干预的重要手段。然而,现有视觉检测算法在复杂解剖环境和光照不佳条件下面临挑战,且声门在操作中尺度变化大——初始为小目标,后迅速占据几乎全部视野。此外,传统方法计算开销高,难以在便携设备上实现实时高精度检测。为此,本文提出一种专为视觉辅助NTI设计的新型声门分割框架。首先,设计轻量级多感受野特征提取模块,降低类内差异,增强尺度鲁棒性,并堆叠形成网络主干与颈部;其次,提出改进的标签分配策略并重定义样本数量,进一步提升复杂环境中分割精度。在三个不同数据集上的实验表明,该网络超越当前最优算法,取得92.9%的分割mDice,模型大小仅19MB,推理速度超过170帧/秒。代码与数据集将开源。

原文摘要 · Abstract (English)

Nasotracheal intubation (NTI) is a critical clinical procedure for establishing and maintaining patient airway patency. Machine-assisted NTI has emerged as a pivotal approach for optimizing procedural efficiency and minimizing manual intervention. However, visual detection algorithms employed for NTI navigation encounter significant challenges, including complex anatomical environments and suboptimal illumination conditions surrounding the glottis. Additionally, the glottis presents considerable scale variability throughout the procedure, initially appearing as a small, difficult-to-capture structure before expanding to occupy nearly the entire field of view. Moreover, traditional visual detection methods often have high computational costs, making real-time, high-precision detection on portable devices challenging. To enhance NTI efficacy and address these challenges, this paper proposes a novel glottis segmentation framework optimized for vision-assisted NTI applications. First, we designed a lightweight, multi-receptive field feature extraction module to reduce intra-class differences, achieving robustness to scale variations of the glottis. This module was then stacked to form the backbone and neck of our network. Subsequently, we developed an advanced label assignment method and redefined the number of samples to further reduce intra-class differences and enhance accuracy in the complex NTI environment. Experiments on three distinct datasets demonstrate that our network surpasses state-of-the-art algorithms, achieving a segmentation mDice of 92.9\% with a compact model size of 19 MB and an inference speed exceeding 170 frames per second. % Our code and datasets will be open-sourced on GitHub after the manuscript is accepted. Our code and datasets are available at https://github.com/HBUT-CV/GlottisNet.

医学图像语义分割实时检测轻量化网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。