arXiv:2506.10653eess.AScs.CL2025-06

用熵最小化和说话人编码,仅用1分钟语音就能稳定适配语音识别器。

Robust Unsupervised Adaptation of a Speech Recogniser Using Entropy Minimisation and Speaker Codes

  • 用多假设条件熵替代单伪标签交叉熵,提升适应鲁棒性。
  • 在Common Voice的远场噪声数据上,1分钟适配使误字率降20%。
  • 适合数据稀缺场景下的语音识别模型快速自适应。

语音识别器通常只在特定环境下表现良好,需适应新环境。针对新说话人适应时数据不足且无标签的问题,本文提出一种结合熵最小化与说话人编码的方法,使仅用1分钟数据的适应仍具鲁棒性。首先,采用基于完整假设的条件熵损失函数,避免依赖单一错误的识别假设;其次,引入短向量形式的“说话人编码”,可低数据量估计。在远场噪声增强版Common Voice数据集上,该方法在1分钟适配数据下实现相对20%的误字率降低,10分钟时提升至29%。

原文摘要 · Abstract (English)

Speech recognisers usually perform optimally only in a specific environment and need to be adapted to work well in another. For adaptation to a new speaker, there is often too little data for fine-tuning to be robust, and that data is usually unlabelled. This paper proposes a combination of approaches to make adaptation to a single minute of data robust. First, instead of estimating the adaptation parameters with cross-entropy on a single error-prone hypothesis or "pseudo-label", this paper proposes a novel loss function, the conditional entropy over complete hypotheses. Using multiple hypotheses makes adaptation more robust to errors in the initial recognition. Second, a "speaker code" characterises a speaker in a vector short enough that it requires little data to estimate. On a far-field noise-augmented version of Common Voice, the proposed scheme yields a 20% relative improvement in word error rate on one minute of adaptation data, increasing on 10 minutes to 29%.

语音识别自适应无监督说话人编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。