arXiv:2606.06806cs.SDeess.AS2026-06中稿 · Interspeech2026

用软分配提升离散语音标记的下游表现,兼顾效率与精度

Leveraging Soft Distributions of SSL-Derived Discrete Speech Tokens for Downstream Inference

论文配图:Leveraging Soft Distributions of SSL-Derived Discrete Speech Tokens for Downstream Inference
图 1 · 摘自论文原文
  • 推理时采用软分配,训练仍用硬分配
  • 在非母语语音识别中超越连续特征模型
  • 生成的表示更贴近音素,泛化能力强

从自监督学习(SSL)模型获得的离散语音标记可实现高效数据压缩并保持强性能,广泛应用于各类任务。但离散化不可避免导致信息损失,使性能低于连续的SSL特征。本文提出仅在下游推理阶段使用软标记分配,训练时仍保持硬分配。该方法在训练阶段保留硬分配的高效性,同时在推理时增强标记表达能力。实验表明,该方法在语音识别(ASR)和语音合成任务中均优于传统硬分配,尤其对域外数据表现出极强泛化能力;在非母语语音识别中甚至超过使用连续SSL特征的模型。对生成表示的分析显示,其与音素对齐更准确。

原文摘要 · Abstract (English)

Discrete speech tokens obtained from self-supervised learning (SSL) models provide efficient data compression while maintaining strong performance, and have been widely used as intermediate representations in various tasks. However, discretization inevitably causes information loss, leading to degraded performance compared with continuous SSL features. In this work, we propose to apply soft token assignment only during downstream inference. This approach preserves the efficiency of hard discretization during training while enhancing the expressiveness of the tokens at inference. The proposed method outperforms conventional hard assignment on both ASR and speech synthesis tasks, and exhibits particularly strong generalizability to out-of-domain data. For ASR of non-native speech, it even surpasses models using continuous SSL features. Moreover, analysis of the resulting representations shows they align more accurately with phonemes compared with conventional hard assignment.

语音识别自监督学习离散表示软分配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。