arXiv:2509.23225cs.CVcs.LG2025-09被引 2

轻量级模型实现实时舌部轮廓分割,适用于多种语言与成像条件。

UltraUNet: Real-Time Ultrasound Tongue Segmentation for Diverse Linguistic and Imaging Conditions

  • 采用轻量化结构与领域优化设计,降低计算与内存开销。
  • 达250帧/秒,单数据集Dice为0.855,跨数据集平均达0.734。
  • 适合语音研究、临床诊断及言语运动障碍分析场景。

超声舌成像(UTI)是一种无创且低成本的研究言语发音、运动控制及相关障碍的工具。然而,由于信噪比低、成像差异大及计算需求高,实时舌轮廓分割仍具挑战。本文提出UltraUNet,一种专为实时舌轮廓分割优化的轻量级编码器-解码器架构。其引入领域特异性创新,包括轻量化Squeeze-and-Excitation模块、小批量稳定性保障的组归一化,以及减少内存开销的求和型跳跃连接,并集成去噪与模糊模拟等超声特定增强。在8个数据集上的评估显示,其表现优异:单数据集Dice为0.855,平均均方距离(MSD)为0.993像素;跨数据集平均Dice为0.734和0.761。UltraUNet为语音研究、临床诊断及言语运动障碍分析提供了快速准确的解决方案。

原文摘要 · Abstract (English)

Ultrasound tongue imaging (UTI) is a non-invasive and cost-effective tool for studying speech articulation, motor control, and related disorders. However, real-time tongue contour segmentation remains challenging due to low signal-to-noise ratios, imaging variability, and computational demands. We propose UltraUNet, a lightweight encoder-decoder architecture optimized for real-time segmentation of tongue contours in ultrasound images. UltraUNet incorporates domain-specific innovations such as lightweight Squeeze-and-Excitation blocks, Group Normalization for small-batch stability, and summation-based skip connections to reduce memory and computational overhead. It achieves 250 frames per second and integrates ultrasound-specific augmentations like denoising and blur simulation. Evaluations on 8 datasets demonstrate high accuracy and robustness, with single-dataset Dice = 0.855 and MSD = 0.993px, and cross-dataset Dice averaging 0.734 and 0.761. UltraUNet provides a fast, accurate solution for speech research, clinical diagnostics, and analysis of speech motor disorders.

超声成像语义分割实时处理语音研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。