arXiv:2603.16920eess.AScs.SD2026-03中稿 · ICASSP 2026被引 1

用大模型生成更真实语音,提升特定领域语音识别准确率

Synthetic Data Domain Adaptation for ASR via LLM-based Text and Phonetic Respelling Augmentation

  • 用大模型生成多样文本并过滤,保证词汇和领域词覆盖
  • 通过伪拼写引入发音变化,合成语音更贴近真实场景
  • 在4个领域数据集上均降低误识率,适合资源少的语音场景

端到端语音识别在特定领域数据上表现下降,主要因域内资源稀缺。本文提出基于合成数据的领域自适应框架,包含两项创新:(1) 基于大语言模型(LLM)的文本增强流程,结合筛选策略,在词汇多样性、困惑度与领域词覆盖率间取得平衡;(2) 一种新型音素重拼增强(PRA)方法,通过大模型生成的正字法伪拼写引入发音多样性。与传统声学级方法(如SpecAugment)不同,PRA在语音合成前即引入发音变异,使合成语音更逼近真实世界的变化。在四个特定领域数据集上的实验表明,该方法持续降低词错误率,验证了结合领域词汇覆盖与真实发音变异性对提升语音识别鲁棒性的显著作用。

原文摘要 · Abstract (English)

End-to-end automatic speech recognition often degrades on domain-specific data due to scarce in-domain resources. We propose a synthetic-data-based domain adaptation framework with two contributions: (1) a large language model (LLM)-based text augmentation pipeline with a filtering strategy that balances lexical diversity, perplexity, and domain-term coverage, and (2) phonetic respelling augmentation (PRA), a novel method that introduces pronunciation variability through LLM-generated orthographic pseudo-spellings. Unlike conventional acoustic-level methods such as SpecAugment, PRA provides phonetic diversity before speech synthesis, enabling synthetic speech to better approximate real-world variability. Experimental results across four domain-specific datasets demonstrate consistent reductions in word error rate, confirming that combining domain-specific lexical coverage with realistic pronunciation variation significantly improves ASR robustness.

语音识别合成数据大模型领域自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。