针对失语和老年语音识别,提出结构化适应方法提升模型泛化能力。
Structured Speaker-Deficiency Adaptation of Foundation Models for Dysarthric and Elderly Speech Recognition
- 分离建模说话人与言语缺陷特征,用独立适配器减少偏差
- 在UASpeech和DementiaBank数据集上分别降低3.01%和1.50%错误率
- 适合低资源失语及老年语音识别,尤其对未见说话人有效
将语音基础模型(SFMs)在稀缺多样的失语和老年语音上进行数据密集型微调,易导致数据偏差且泛化能力差。本文提出新型结构化说话人缺陷自适应方法,用于自监督预训练的SFMs。在有监督微调阶段构建对说话人和言语缺陷不变的SFMs,减少对训练说话人的过度依赖,为测试时无监督自适应提供更中立稳健的起点。通过独立适配器分别建模说话人身份、言语障碍严重程度或老龄化引起的神经认知衰退带来的语音差异,并可组合以建模任意已见或未见说话人。在UASpeech失语语音和DementiaBank Pitt老年语音数据集上的实验表明,基于HuBERT和Wav2vec2-conformer的结构化自适应方法,显著优于基线模型(无适配器、全局共享适配器、单属性适配器),在两项任务上绝对词错误率分别降低3.01%和1.50%(相对降低10.86%和6.94%)。在包含16名失语者的UASpeech测试集上取得最低公开报道的词错误率19.45%(极低可懂度下49.34%,未见词33.17%)。
原文摘要 · Abstract (English)
Data-intensive fine-tuning of speech foundation models (SFMs) to scarce and diverse dysarthric and elderly speech leads to data bias and poor generalization to unseen speakers. This paper proposes novel structured speaker-deficiency adaptation approaches for SSL pre-trained SFMs on such data. Speaker and speech deficiency invariant SFMs were constructed in their supervised adaptive fine-tuning stage to reduce undue bias to training data speakers, and serves as a more neutral and robust starting point for test time unsupervised adaptation. Speech variability attributed to speaker identity and speech impairment severity, or aging induced neurocognitive decline, are modelled using separate adapters that can be combined together to model any seen or unseen speaker. Experiments on the UASpeech dysarthric and DementiaBank Pitt elderly speech corpora suggest structured speaker-deficiency adaptation of HuBERT and Wav2vec2-conformer models consistently outperforms baseline SFMs using either: a) no adapters; b) global adapters shared among all speakers; or c) single attribute adapters modelling speaker or deficiency labels alone by statistically significant WER reductions up to 3.01% and 1.50% absolute (10.86% and 6.94% relative) on the two tasks respectively. The lowest published WER of 19.45% (49.34% on very low intelligibility, 33.17% on unseen words) is obtained on the UASpeech test set of 16 dysarthric speakers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。