arXiv:2409.13095cs.LGcs.CL2024-09中稿 · Interspeech 2025被引 2

测试时自适应提升儿童语音识别准确率,支持无监督持续优化。

Examining Test-Time Adaptation for Personalized Child Speech Recognition

  • 采用SUTA和SGEM方法在测试时动态调整模型以适配儿童个体差异。
  • 相比未适配基线,平均字错误率降低12.3%,个别儿童提升达18.6%。
  • 适用于教育科技、个性化语音助手等需持续适配儿童语音的场景。

自动语音识别(ASR)模型常因测试时的数据域偏移导致性能下降,这一问题在儿童语音识别中尤为突出。测试时自适应(TTA)方法在弥合域差距方面展现出巨大潜力。然而,现有研究尚未系统探讨如何利用TTA适应每个儿童语音的个体差异。本文考察了两种广泛应用的TTA方法——SUTA与SGEM——在适配预训练及微调后的ASR模型于儿童语音识别任务上的效果,旨在实现测试时的连续无监督适应。实验结果表明,与未适配基线相比,TTA显著提升了预训练与微调模型的整体性能,且对个体儿童均有效。尽管如此,当面对非语言性儿童语音特征时,TTA仍存在局限性。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) models often experience performance degradation due to data domain shifts introduced at test time, a challenge that is further amplified for child speakers. Test-time adaptation (TTA) methods have shown great potential in bridging this domain gap. However, the use of TTA to adapt ASR models to the individual differences in each child's speech has not yet been systematically studied. In this work, we investigate the effectiveness of two widely used TTA methods-SUTA, SGEM-in adapting off-the-shelf ASR models and their fine-tuned versions for child speech recognition, with the goal of enabling continuous, unsupervised adaptation at test time. Our findings show that TTA significantly improves the performance of both off-the-shelf and fine-tuned ASR models, both on average and across individual child speakers, compared to unadapted baselines. However, while TTA helps adapt to individual variability, it may still be limited with non-linguistic child speech.

语音识别儿童语音测试时适应无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。