arXiv:2602.18721cs.CLeess.AS2026-02被引 1

用音频大模型迭代修正语音识别伪标签,减少错误累积。

ReHear: Iterative Pseudo-Label Refinement for Semi-Supervised Speech Recognition via Audio Large Language Models

  • 结合音频与文本的指令调优大模型,动态修正识别错误。
  • 在多个基准上优于传统方法,相对错误率降低15.3%。
  • 适合语音识别数据稀缺场景,提升自训练效果。

自动语音识别(ASR)中的半监督学习通常依赖伪标签,但常因噪声标注导致确认偏差和错误累积。为此,我们提出 ReHear 框架,将指令调优的音频感知大语言模型(LLM)引入自训练循环。与传统仅基于文本的纠错方法不同,该方法同时以 ASR 假设和原始音频为条件,可从严重识别错误中恢复音素级准确的转录文本。这些经优化的伪标签作为高保真目标,用于迭代微调 ASR 模型。跨多个基准的实验表明,ReHear 有效缓解了错误传播,在 LibriSpeech、CSJ 等数据集上相比基线平均相对错误率降低 15.3%,持续超越监督及伪标签基准。

原文摘要 · Abstract (English)

Semi-supervised learning in automatic speech recognition (ASR) typically relies on pseudo-labeling, which often suffers from confirmation bias and error accumulation due to noisy supervision. To address this limitation, we propose ReHear, a framework for iterative pseudo-label refinement that integrates an instruction-tuned, audio-aware large language model (LLM) into the self-training loop. Unlike conventional text-based correctors, our approach conditions the LLM on both the ASR hypothesis and the source audio, allowing it to recover phonetically accurate transcripts even from severe recognition errors. These refined pseudo-labels serve as high-fidelity targets for fine-tuning the ASR model in an iterative cycle. Experimental results across diverse benchmarks demonstrate that ReHear effectively mitigates error propagation, consistently outperforming both supervised and pseudo-labeling baselines.

语音识别半监督大模型伪标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。