arXiv:2506.04586cs.CLcs.SD2025-06中稿 · ICASSP 2026被引 1

用大模型修正语音数据伪标签,提升真实场景下的语音模型性能。

LESS: Large Language Model Enhanced Semi-Supervised Learning for Speech Foundational Models Using in-the-wild Data

  • 用大模型自动修正真实环境语音的伪标签
  • 在中文语音识别上降低3.8%错误率,翻译任务提升0.7~0.8分
  • 适合做语音基础模型训练,尤其适用于复杂真实场景

尽管先进的语音基础模型能生成高质量的文本伪标签,但在真实世界数据上应用半监督学习仍具挑战,因其声学环境更丰富复杂。为此,我们提出LESS(大语言模型增强的半监督学习),利用大语言模型对真实环境数据生成的伪标签进行修正,并结合数据过滤策略进一步优化。在中文语音识别和西班牙语到英语语音翻译任务中,LESS均表现稳定:在WenetSpeech数据集上实现3.8%的绝对词错误率降低;在Callhome和Fisher测试集上分别获得34.0和64.7的得分,BLEU值提升0.8和0.7。结果表明LESS在多种语言、任务和领域中均具有效性。相关训练方案已开源,以促进该方向研究。

原文摘要 · Abstract (English)

Although state-of-the-art Speech Foundational Models can produce high-quality text pseudo-labels, applying Semi-Supervised Learning (SSL) for in-the-wild real-world data remains challenging due to its richer and more complex acoustics compared to curated datasets. To address the challenges, we introduce LESS (Large Language Model Enhanced Semi-supervised Learning), a versatile framework that uses Large Language Models (LLMs) to correct pseudo-labels generated on in-the-wild data. In the LESS framework, pseudo-labeled text from Automatic Speech Recognition (ASR) or Automatic Speech Translation (AST) of the unsupervised data is refined by an LLM, and further improved by a data filtering strategy. Across Mandarin ASR and Spanish-to-English AST evaluations, LESS delivers consistent gains, with an absolute Word Error Rate reduction of 3.8% on WenetSpeech, and BLEU score increase of 0.8 and 0.7, achieving 34.0 on Callhome and 64.7 on Fisher testsets respectively. These results highlight LESS's effectiveness across diverse languages, tasks, and domains. We have released the recipe as open source to facilitate further research in this area.

语音识别半监督学习大模型应用真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。