用语音合成生成老人说话数据,显著降低老年语音识别错误率。
Elderly-Contextual Data Augmentation via Speech Synthesis for Elderly ASR

- 用大模型改写语料,再通过老人音色合成新语音数据。
- 在英韩语老年语音集上,错误率最高降低58.2%。
- 适合资源少的老年语音识别研究者使用。
尽管自动语音识别(ASR)取得进展,老年语音识别(EASR)仍因训练数据稀缺及老年人语音特有的声学与语言特征而困难。本文提出一种结合大语言模型(LLM)文本改写与语音合成(TTS)的数据增强流程:给定老年语音数据集,先由LLM生成具有老年人语境的语料改写版本,再用老年参考发音人合成对应语音。合成的音文对与原始数据合并,用于微调Whisper模型,无需修改结构。进一步分析了低资源场景下增强比例与参考发音人构成的影响。在70岁以上说话者的英、韩语老年语音数据集上实验表明,该方法持续优于传统增强基线,相比Whisper基线最高实现58.2%的词错误率(WER)下降。
原文摘要 · Abstract (English)
Despite recent progress in automatic speech recognition (ASR), elderly ASR (EASR) remains challenging due to limited training data and the distinct acoustic and linguistic characteristics of elderly speech. In this work, we address data scarcity in EASR through a data augmentation pipeline that combines large language model (LLM)-based transcript paraphrasing with text-to-speech (TTS) synthesis. Given an elderly speech dataset, the LLM first generates elderly-contextual paraphrases of the original transcripts, and the TTS model then synthesizes corresponding speech using elderly reference speakers. The resulting synthetic audio-text pairs are merged with the original data to fine-tune Whisper without architectural modification. We further analyze the effects of augmentation ratio and reference-speaker composition in low-resource EASR. Experiments on English and Korean elderly speech datasets from speakers aged 70 and above show that the proposed method consistently improves performance over conventional augmentation baselines, achieving up to a 58.2% reduction in word error rate (WER) compared with the Whisper baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。