arXiv:2608.00186eess.SPcs.SD2026-08

用正常语音做锚点,让语音识别模型更好听懂唇腭裂患者的说话。

Normal-Anchored First-Order Model-Agnostic Meta-Learning based Whisper Fine-Tuning for Enhancing Fairness of Cleft Lip and Palate Speech Recognition

论文配图:Normal-Anchored First-Order Model-Agnostic Meta-Learning based Whisper Fine-Tuning for Enhancing Fairness of Cleft Lip and Palate Speech Recognition
图 1 · 摘自论文原文
  • 以正常语音为稳定锚点,在元学习中提升对不同严重程度唇腭裂语音的适应能力。
  • 在NMCPC和AIISH数据集上,严重病例的误识率仍高达52%和57.5%,但整体鲁棒性提升。
  • 适合需要提升罕见病语音识别公平性的研究者和医疗人工智能开发者。

唇腭裂(CLP)语音的自动语音识别(ASR)难度大,因发音特征随严重程度变化显著,导致预训练系统性能下降,常规微调在低资源、异质条件下泛化能力差。本文提出基于正常锚定的一阶模型无关元学习(NA-FOMAML)方法,将正常语音用于内层循环作为稳定支持,唇腭裂严重程度分组用于外层循环以增强适应后鲁棒性。在NMCPC和AIISH数据集上,采用四种正常锚定训练配置:冻结编码器、全编码器及不同层段(0-5, 4-11, 6-11, 8-11)微调策略,同时适配解码器与投影头。结果表明,仅用正常语音进行外层训练效果不足。在NMCPC上,全编码器微调从正常到正常+轻度+中度时,正常、轻度、中度、重度语音的词错误率(WER)分别为4.40%、5.53%、16.14%、52.07%;在AIISH上,从正常到正常+轻度+中度+重度,对应为2.48%、19.66%、14.05%、57.50%。基于音素类别的转录分析显示,重度患者在擦音、塞擦音、鼻音、流音、塞音和元音上均存在高错误率。总体而言,NA-FOMAML提升了跨严重程度的鲁棒性,但重度语音仍需针对性采样、音素感知损失函数及针对压力辅音和共振畸变的增强策略。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) for cleft lip and palate (CLP) speech is difficult because acoustic and articulatory patterns vary across severity levels. This variability reduces the performance of pretrained ASR systems, and conventional fine-tuning may not generalize well under low-resource, heterogeneous CLP conditions. This work proposes Normal-Anchored First-Order Model-Agnostic Meta-Learning (NA-FOMAML) for adapting Whisper to CLP speech. The method uses a first-order bilevel meta-learning framework in which normal speech is used in the inner loop as a stable support condition, while CLP severity groups are used in the outer loop to improve post-adaptation robustness. This design aims to reduce the performance gap between normal and pathological speech. Experiments are conducted on the NMCPC and AIISH datasets using four normal-anchored training configurations. Frozen encoder, full encoder, and selected Whisper encoder-layer tuning strategies are evaluated, including layers 0--5, 4--11, 6--11, and 8--11, with decoder and projection-head adaptation. Results show that outer-loop training with only normal speech is insufficient. For NMCPC, full encoder tuning with Normal to Normal+Mild+Moderate gives WERs of 4.40%, 5.53%, 16.14%, and 52.07% for normal, mild, moderate, and severe speech. For AIISH, full encoder tuning with Normal to Normal+Mild+Moderate+Severe gives WERs of 2.48%, 19.66%, 14.05%, and 57.50%. A transcription-based phoneme-category analysis shows that severe CLP speech has high error rates across fricatives, affricates, nasals, liquids, plosives, and vowels. Overall, NA-FOMAML improves cross-severity robustness, but severe speech still requires severity-aware sampling, phoneme-aware loss functions, and augmentation targeting pressure consonant and resonance-related distortions.

语音识别唇腭裂元学习公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。