arXiv:2603.02285cs.SDcs.LG2026-03中稿 · ICASSP 2026

提出一种新损失函数,让语音识别在无配对数据下也能有效训练。

Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study

  • 基于分类误差界构建理论框架,揭示无监督语音识别的可行条件。
  • 推导出可验证的误差上界,证明方法在特定条件下有效。
  • 设计单阶段序列级交叉熵损失,适合无标注语音数据训练。

无监督语音识别是仅用未配对数据训练语音识别模型的任务。为明确无监督语音识别何时及如何成功,以及分类误差与候选训练目标的关系,本文建立了一个基于分类误差界的理论框架。提出了两个无监督语音识别可行的条件,并讨论了其必要性。在这些条件下,推导出无监督语音识别的分类误差上界,并通过模拟验证了该上界。受此上界启发,提出一种单阶段序列级交叉熵损失,用于无监督语音识别。

原文摘要 · Abstract (English)

Unsupervised speech recognition is a task of training a speech recognition model with unpaired data. To determine when and how unsupervised speech recognition can succeed, and how classification error relates to candidate training objectives, we develop a theoretical framework for unsupervised speech recognition grounded in classification error bounds. We introduce two conditions under which unsupervised speech recognition is possible. The necessity of these conditions are also discussed. Under these conditions, we derive a classification error bound for unsupervised speech recognition and validate this bound in simulations. Motivated by this bound, we propose a single-stage sequence-level cross-entropy loss for unsupervised speech recognition.

语音识别无监督学习理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。