arXiv:2505.16207cs.SDeess.AS2025-05中稿 · Interspeech2025被引 7

用可微分k-means优化语音识别的离散编码,提升识别精度。

Differentiable K-means for Fully-optimized Discrete Token-based ASR

  • 提出可微分k-means,联合优化语音编码与下游任务。
  • 在ASR任务中显著提升准确率,离散码更具音素纯度。
  • 适合需要高质量语音表示的研究者,如语音识别与合成。

近期研究显示,自监督学习(SSL)模型生成的离散标记在多种语音任务中具有潜力。这些标记不仅可替代文本进行语言建模,还可作为自动语音识别(ASR)等任务的中间表示。然而,传统k-means聚类独立于下游任务,导致标记非最优。本文提出可微分k-means,实现标记化与下游任务的联合优化,支持对SSL参数及多层输出权重的微调。以ASR为下游任务进行实验,结果表明优化后的标记显著提升了识别准确率,且其音素信息纯度更高,甚至在语音重合成中也表现出良好效果。

原文摘要 · Abstract (English)

Recent studies have highlighted the potential of discrete tokens derived from self-supervised learning (SSL) models for various speech-related tasks. These tokens serve not only as substitutes for text in language modeling but also as intermediate representations for tasks such as automatic speech recognition (ASR). However, discrete tokens are typically obtained via k-means clustering of SSL features independently of downstream tasks, making them suboptimal for specific applications. This paper proposes the use of differentiable k-means, enabling the joint optimization of tokenization and downstream tasks. This approach enables the fine-tuning of the SSL parameters and learning weights for outputs from multiple SSL layers. Experiments were conducted with ASR as a downstream task. ASR accuracy successfully improved owing to the optimized tokens. The acquired tokens also exhibited greater purity of phonetic information, which were found to be useful even in speech resynthesis.

语音识别自监督学习可微分聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。