arXiv:2604.27273cs.SD2026-04中稿 · ICML

用合成口音语音提升低资源ASR,随机音素替换比大模型生成更有效。

Few-Shot Synthetic Accented Speech for ASR Fine-Tuning: What Helps and When?

论文配图:Few-Shot Synthetic Accented Speech for ASR Fine-Tuning: What Helps and When?
图 1 · 摘自论文原文
  • 用随机音素替换模拟口音,效果接近真实口音数据。
  • 合成数据量增大时,真实口音数据优势减弱,比例关键。
  • 混合真实与合成数据可稳定训练,但过量会稀释真实信息。

当真实口音语音稀缺时,合成口音语音是提升自动语音识别(ASR)性能的可行方案。本文探究何种合成方式更有效:目标口音音素编辑暴露识别器于特定发音模式,或随机音素扰动作为音素空间中的数据增强。在少样本文本转语音(TTS)流程中,对比了大语言模型生成的口音编辑、等量随机替换及基于真实口音音素和语调的基准控制实验。结果显示,随机替换已能实现大部分ASR性能提升:大模型生成的口音编辑仅小幅优于随机替换,真实口音音素表现接近随机基线,且随合成数据集扩大趋于收敛;加入真实语调仅带来小幅增益。混合真实与合成语音可稳定低资源微调,但固定合成预算后期会稀释真实数据信息,表明真实-合成比例至关重要。

原文摘要 · Abstract (English)

Synthetic accented speech is a promising way to improve automatic speech recognition (ASR) when real accented recordings are scarce. We ask what makes such data useful for ASR fine-tuning: target-accent phoneme edits that expose the recognizer to accent-specific pronunciations, or random phoneme perturbations that act as augmentation in phoneme space. In a few-shot TTS pipeline, we compare LLM-generated accent edits with matched-rate random substitutions and oracle controls using ground-truth accented phonemes and prosody. Random substitutions recover much of the ASR gain: LLM target-accent edits improve over random by only a small margin, ground-truth phonemes stay close to the random baseline and nearly converge with it as the synthetic ASR fine-tuning set grows larger, and adding ground-truth prosody yields only a modest further gain. Mixing synthetic with real accented speech also stabilizes low-resource fine-tuning, but a fixed synthetic budget can later dilute the information in real data, showing that the real--synthetic ratio matters.

语音识别合成数据少样本学习口音适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。