arXiv:2501.09113eess.AScs.SD2025-01被引 1

针对用户音频特征定制数据增强,提升个性化语音识别准确率

persoDA: Personalized Data Augmentation for Personalized ASR

  • 根据用户数据定制增强策略,而非随机加噪混响
  • 相比标准增强方法,相对降低13.9%的词错误率
  • 训练收敛速度提升16%至20%,适合移动端个性化模型

数据增强(DA)广泛应用于自动语音识别(ASR)模型训练中,可提升数据多样性、鲁棒性与泛化能力。近期研究表明,在移动设备上对ASR模型进行个性化可降低词错误率(WER)。本文评估了该场景下的数据增强方法,并提出persoDA:一种基于用户数据驱动的个性化数据增强方法,旨在针对终端用户的声学特征进行训练数据增强,区别于传统的多条件训练(MCT)中随机混入混响和噪声的方式。在基于Conformer架构、于LibriSpeech上训练并针对VOICES数据集进行个性化的实验中,persoDA相较标准数据增强(随机噪声与混响)实现13.9%的相对WER降低;同时,训练收敛速度比MCT快16%至20%。

原文摘要 · Abstract (English)

Data augmentation (DA) is ubiquitously used in training of Automatic Speech Recognition (ASR) models. DA offers increased data variability, robustness and generalization against different acoustic distortions. Recently, personalization of ASR models on mobile devices has been shown to improve Word Error Rate (WER). This paper evaluates data augmentation in this context and proposes persoDA; a DA method driven by user's data utilized to personalize ASR. persoDA aims to augment training with data specifically tuned towards acoustic characteristics of the end-user, as opposed to standard augmentation based on Multi-Condition Training (MCT) that applies random reverberation and noises. Our evaluation with an ASR conformer-based baseline trained on Librispeech and personalized for VOICES shows that persoDA achieves a 13.9% relative WER reduction over using standard data augmentation (using random noise & reverberation). Furthermore, persoDA shows 16% to 20% faster convergence over MCT.

语音识别个性化数据增强移动端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。