arXiv:2509.15373cs.CL2025-09被引 3

用生成文本+语音合成,低成本提升低资源语音识别效果

Frustratingly Easy Data Augmentation for Low-Resource ASR

  • 用词汇替换或大模型生成新文本,再转成合成语音
  • 在4种低资源语言上,最高降低14.3%的错误率
  • 方法简单通用,适合数据少的场景

本文提出三种自包含的低资源语音识别数据增强方法。先通过基于释义词替换、随机替换或大模型生成新文本,再用文语转换(TTS)生成合成语音。仅使用原始标注数据,在四种资源极度匮乏的语言(Vatlongos、Nashta、Shinekhen Buryat、Kakabe)上进行实验。将预训练的Wav2Vec2-XLSR-53模型在原始音频与合成数据组合上微调,显著提升性能,其中Nashta语言的词错误率(WER)绝对降低14.3%。该方法在所有四类低资源语言中均有效,并在英语等高资源语言上也表现良好,证明其广泛适用性。

原文摘要 · Abstract (English)

This paper introduces three self-contained data augmentation methods for low-resource Automatic Speech Recognition (ASR). Our techniques first generate novel text--using gloss-based replacement, random replacement, or an LLM-based approach--and then apply Text-to-Speech (TTS) to produce synthetic audio. We apply these methods, which leverage only the original annotated data, to four languages with extremely limited resources (Vatlongos, Nashta, Shinekhen Buryat, and Kakabe). Fine-tuning a pretrained Wav2Vec2-XLSR-53 model on a combination of the original audio and generated synthetic data yields significant performance gains, including a 14.3% absolute WER reduction for Nashta. The methods prove effective across all four low-resource languages and also show utility for high-resource languages like English, demonstrating their broad applicability.

语音识别数据增强低资源TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。