用拼音级语音合成提升语音识别,关键在选对文本和参考音频。
Scaling phoneme-based TTS augmentation for ASR: A unified pipeline and controlled study
- 基于拼音的统一语音合成流程,支持多语言自动构建训练数据。
- 按发音频率选句子,比随机选能降低19.3%错误率。
- 筛选高质量参考语音,让法语、意大利语识别准确率显著提升。
合成语音可为自动语音识别(ASR)提供可扩展的监督信号,但其效果取决于所选文本、参考语音及合成数据量。本文提出一个基于拼音的统一TTS-to-ASR增强流水线,使用从头训练的多语言TTS模型(F5-TTS架构,带语言标识条件),结合语言特定的音素转写、参考语音筛选、候选文本选择、合成与匹配的ASR续接。我们进一步提出音素频率引导选择(PFGS),根据真实ASR训练标签估算音素频率来排序候选句子。在阿拉伯语、法语、意大利语和葡萄牙语的独立单语ASR系统上,共覆盖13个测试集。在合成规模测试中,随机增强在11个测试集上优于仅用真实数据的延续。在60%合成预算下,PFGS在12个测试集上优于纯真实训练,在9个测试集上优于随机选择,最大相对词错误率(WER)降低19.3%。固定目标文本与合成数量时,参考语音筛选使意大利语和法语Common Voice数据集的绝对WER分别降低0.29和0.59点。结果表明,合成规模、候选文本内容与参考语音质量是TTS增强中的关键控制变量。
原文摘要 · Abstract (English)
Synthetic speech provides scalable supervision for automatic speech recognition (ASR), but its benefit depends on the selected texts, reference speech, and amount of synthesized data. We present a unified phoneme-based TTS-to-ASR augmentation pipeline built around a multilingual TTS model trained from scratch using the F5-TTS architecture with language-ID conditioning. The pipeline combines language-specific grapheme-to-phoneme conversion, reference-speech filtering, candidate-text selection, synthesis, and matched ASR continuation. We further propose phoneme-frequency-guided selection (PFGS), which ranks candidate sentences using phoneme frequencies estimated from real ASR training labels. Experiments with separate monolingual ASR systems for Arabic, French, Italian, and Portuguese span 13 test sets. Across the synthesis-scale sweep, random augmentation improves over matched real-only continuation on 11 test sets. Under a nominal 60% synthesis budget, PFGS improves over real-only training on 12 test sets and over random selection on 9. Its largest relative word error rate (WER) reduction against random selection is 19.3%. With target texts and synthesis counts fixed, reference-speech filtering reduces absolute WER by 0.29 and 0.59 points on Italian and French Common Voice, respectively. These results identify synthesis scale, candidate-text content, and reference quality as important control variables in TTS-based ASR augmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。