用大模型生成对话数据,让低资源语音识别更准。
Efficient ASR Training with Conversations that Never Happened

- 用大模型生成带角色信息的虚拟对话,再合成语音。
- 67小时真实+636小时模拟数据胜过2700小时真实数据。
- 适合低资源语言或小众领域的语音识别训练。
针对低资源语言和特定领域中的对话式语音识别,受限于匹配的多说话人训练数据稀缺。本文提出一种增强流水线:生成带参与者元信息的场景级对话,将说话人属性映射至文本转语音(TTS)声线,再将合成语句组装成具备说话人感知的模拟对话。在匈牙利语BEA-Dialogue基准上,使用相同FastConformer-Large训练配方,评估了五种大语言模型家族在单生成器、固定预算混合及扩展规模设置下的表现。结果表明,合成对话能持续提升语音识别性能,但生成器选择与数据构成显著影响增益。最大配置仅需67小时真实对话与636小时模拟数据,即优于零样本模型在2700小时匈牙利语数据上的表现。研究证实,结合大模型与TTS生成的对话数据,可有效补充真实语料用于语音模型训练。
原文摘要 · Abstract (English)
Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations. We evaluated five LLM families under single-generator, fixed-budget mixture, and scale-up settings using the same FastConformer-Large training recipe for each one. We ran comprehensive evaluations on the Hungarian BEA-Dialogue benchmark corpus, with the method itself being applicable to any language given the resources for each component. The results show that synthetic conversations consistently improve speech recognition performance, but generator choice and data composition strongly affect the gains. Our largest training configuration, using only 67 hours of real conversations and 636 hours of simulated data, achieves better performance on the evaluation benchmark than a zero-shot model trained on 2700 hours of Hungarian speech. These findings indicate that LLM-generated conversational data synthesized with TTS is a practical complement to real conversational corpora for speech model training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。