arXiv:2411.18320cs.CLcs.AI2024-11被引 1

用语音链框架+梯度记忆,实现语音识别持续学习不遗忘。

Continual Learning in Machine Speech Chain Using Gradient Episodic Memory

  • 构建语音链框架,用语音合成回放旧任务数据
  • 在LJ Speech上误差率显著降低,噪声下性能稳定
  • 适合需持续更新的语音识别系统开发者

自动语音识别(ASR)系统的持续学习面临挑战,尤其在避免灾难性遗忘的同时保持对先前任务的性能。本文提出一种新方法,利用机器语音链框架结合梯度回忆记忆(GEM),在ASR中实现持续学习。通过在语音链中引入文本转语音(TTS)组件,支持GEM所需的回放机制,使ASR模型能顺序学习新任务而不明显损害旧任务表现。在LJ Speech数据集上的实验表明,该方法优于传统微调和多任务学习,在不同噪声条件下均实现显著误差率下降,展现了半监督语音链在语音识别持续学习中的有效性与高效性。

原文摘要 · Abstract (English)

Continual learning for automatic speech recognition (ASR) systems poses a challenge, especially with the need to avoid catastrophic forgetting while maintaining performance on previously learned tasks. This paper introduces a novel approach leveraging the machine speech chain framework to enable continual learning in ASR using gradient episodic memory (GEM). By incorporating a text-to-speech (TTS) component within the machine speech chain, we support the replay mechanism essential for GEM, allowing the ASR model to learn new tasks sequentially without significant performance degradation on earlier tasks. Our experiments, conducted on the LJ Speech dataset, demonstrate that our method outperforms traditional fine-tuning and multitask learning approaches, achieving a substantial error rate reduction while maintaining high performance across varying noise conditions. We showed the potential of our semi-supervised machine speech chain approach for effective and efficient continual learning in speech recognition.

持续学习语音识别语音链GEM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。