arXiv:2506.11089eess.AScs.AI2025-06被引 9

用大模型统一融合多语音识别结果,提升伪标签准确率

Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM

  • 用大语言模型统一处理多个ASR输出,替代传统投票机制
  • 在多个数据集上,伪标签准确率显著高于传统方法
  • 适合做自监督语音识别的科研人员和工程师

自动语音识别(ASR)模型依赖高质量转录数据进行训练。为大规模无标签音频数据生成伪标签时,传统方法常采用多阶段处理的复杂流程,导致误差传播、信息丢失和优化不一致。本文提出一种统一的多ASR提示驱动框架,通过文本或语音类大语言模型(LLM)进行后处理,取代传统的投票等仲裁逻辑以整合集成结果。我们对比了多种带与不带LLM的架构,发现使用LLM的方案在转录准确率上显著优于传统方法。此外,利用不同方法生成的伪标签训练半监督ASR模型,在多个数据集上均显示基于文本和语音LLM的转录结果带来更好的性能表现。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) models rely on high-quality transcribed data for effective training. Generating pseudo-labels for large unlabeled audio datasets often relies on complex pipelines that combine multiple ASR outputs through multi-stage processing, leading to error propagation, information loss and disjoint optimization. We propose a unified multi-ASR prompt-driven framework using postprocessing by either textual or speech-based large language models (LLMs), replacing voting or other arbitration logic for reconciling the ensemble outputs. We perform a comparative study of multiple architectures with and without LLMs, showing significant improvements in transcription accuracy compared to traditional methods. Furthermore, we use the pseudo-labels generated by the various approaches to train semi-supervised ASR models for different datasets, again showing improved performance with textual and speechLLM transcriptions compared to baselines.

语音识别伪标签大语言模型半监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。