arXiv:2601.00100eess.AScs.CL2026-01Transactions of th…

用变分预测编码解释HuBERT,提升语音表征学习效果

Learning Speech Representations with Variational Predictive Coding

  • 从变分视角揭示HuBERT的预测编码本质
  • 两项简单修改使预训练性能显著提升
  • 适用于语音识别、声调追踪等下游任务

尽管HuBERT是当前最知名的语音表示学习目标,但其发展已停滞。本文指出,其缺乏底层原理是瓶颈,并证明在变分视角下的预测编码正是HuBERT的核心原理。该统一框架为参数化与优化提供新思路,我们提出两种简单改进,即刻提升性能。此外,该框架与APC、CPC、wav2vec、BEST-RQ等目标存在紧密联系。实验表明,预训练改进显著提升四项下游任务表现:音素分类、基频追踪、说话人识别和自动语音识别,凸显预测编码解释的重要性。

原文摘要 · Abstract (English)

Despite being the best known objective for learning speech representations, the HuBERT objective has not been further developed and improved. We argue that it is the lack of an underlying principle that stalls the development, and, in this paper, we show that predictive coding under a variational view is the principle behind the HuBERT objective. Due to its generality, our formulation provides opportunities to improve parameterization and optimization, and we show two simple modifications that bring immediate improvements to the HuBERT objective. In addition, the predictive coding formulation has tight connections to various other objectives, such as APC, CPC, wav2vec, and BEST-RQ. Empirically, the improvement in pre-training brings significant improvements to four downstream tasks: phone classification, f0 tracking, speaker recognition, and automatic speech recognition, highlighting the importance of the predictive coding interpretation.

语音表征预测编码自监督学习HuBERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。