arXiv:2603.17383eess.AS2026-03

通过聚焦鼻音特征学习,提升语音筛查在真实环境下的鲁棒性。

Robust Nasality Representation Learning for Cleft Palate-Related Velopharyngeal Dysfunction Screening in Real-World Settings

  • 先用语音对比预训练学习鼻音相关表征,再用轻量分类器判断
  • 跨域测试中准确率达69.5%,优于现有最佳基线(64.1%)
  • 适合临床部署,对设备、噪声等差异不敏感

腭裂相关的咽部闭合不全(VPD)表现为发音时软腭无法完全闭合,常导致过度鼻音和可懂度下降。尽管基于语音的机器学习模型在标准化临床录音条件下表现良好,但在真实场景中因设备、通道、噪声和声学环境差异导致的领域偏移,性能常大幅下降。为此,我们提出一种两阶段框架用于VPD筛查:首先,在带音素对齐的辅助语料上,通过口部上下文与鼻部上下文的监督信号,进行有监督对比预训练,学习聚焦鼻音的语音表征;其次,冻结编码器,使用轻量级分类器处理0.5秒语音片段,汇总概率并以固定阈值输出录音级决策。在82名受试者的域内临床队列中,该方法达到完美筛查性能(宏平均F1=1.000,准确率=1.000)。在包含131个异构公开网络录音的域外数据集上,大模型预训练表示性能显著下降,而MFCC为最强基线(宏平均F1=0.612,准确率=0.641)。所提方法在域外表现最优(宏平均F1=0.679,准确率=0.695),在相同评估协议下超越最强基线。结果表明,预先学习聚焦鼻音的表征可降低对录音伪影的敏感性,提升可部署语音筛查的鲁棒性。

原文摘要 · Abstract (English)

Velopharyngeal dysfunction (VPD) is characterized by inadequate velopharyngeal closure during speech and often causes hypernasality and reduced intelligibility. Although speech-based machine learning models can perform well under standardized clinical recording conditions, their performance often drops in real-world settings because of domain shift caused by differences in devices, channels, noise, and room acoustics. To improve robustness, we propose a two-stage framework for VPD screening. First, a nasality-focused speech representation is learned by supervised contrastive pre-training on an auxiliary corpus with phoneme alignments, using oral-context versus nasal-context supervision. Second, the encoder is frozen and used with lightweight classifiers on 0.5-second speech chunks, whose probabilities are aggregated to produce recording-level decisions with a fixed threshold. On an in-domain clinical cohort of 82 subjects, the proposed method achieved perfect recording-level screening performance (macro-F1 = 1.000, accuracy = 1.000). On a separate out-of-domain set of 131 heterogeneous public Internet recordings, large pretrained speech representations degraded substantially, while MFCC was the strongest baseline (macro-F1 = 0.612, accuracy = 0.641). The proposed method achieved the best out-of-domain performance (macro-F1 = 0.679, accuracy = 0.695), improving on the strongest baseline under the same evaluation protocol. These results suggest that learning a nasality-focused representation before clinical classification can reduce sensitivity to recording artifacts and improve robustness for deployable speech-based VPD screening.

语音筛查鼻音表征鲁棒性临床应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。