arXiv:2505.20216eess.AS2025-05中稿 · INTERSPEECH 2025被引 1

解决儿童语音识别持续学习中的遗忘问题,提升隐私保护下的模型适应性。

Continuous Learning for Children's ASR: Overcoming Catastrophic Forgetting with Elastic Weight Consolidation and Synaptic Intelligence

  • 采用EWC与SI技术缓解在线学习时的灾难性遗忘。
  • 在MyST语料上实现相对词错误率降低5.21%(EWC)和4.36%(SI)。
  • 适用于需持续更新、注重儿童隐私的语音识别场景。

本文首次研究了儿童语音识别在在线学习场景下的持续学习问题,这对面向儿童的应用及未成年人隐私保护具有重要意义。传统微调方法常因灾难性遗忘导致性能下降。我们探索了弹性权重巩固(EWC)与突触智能(SI)两种成熟技术,在针对在线学习定制的MyST语料协议下,相较微调基线,分别实现5.21%和4.36%的相对词错误率(WER)降低。

原文摘要 · Abstract (English)

In this work, we present the first study addressing automatic speech recognition (ASR) for children in an online learning setting. This is particularly important for both child-centric applications and the privacy protection of minors, where training models with sequentially arriving data is critical. The conventional approach of model fine-tuning often suffers from catastrophic forgetting. To tackle this issue, we explore two established techniques: elastic weight consolidation (EWC) and synaptic intelligence (SI). Using a custom protocol on the MyST corpus, tailored to the online learning setting, we achieve relative word error rate (WER) reductions of 5.21% with EWC and 4.36% with SI, compared to the fine-tuning baseline.

语音识别持续学习儿童语音隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。