arXiv:2512.03301eess.AS2025-12被引 2

对比监督与无监督语义语音标记,发现监督方法在儿童语音识别中更优。

Comparing Unsupervised and Supervised Semantic Speech Tokens: A Case Study of Child ASR

  • 用监督学习训练语义语音标记,提升模型性能。
  • 监督方法在极低码率下仍优于连续表示和无监督方法。
  • 适合低资源场景的语音识别研究者参考。

离散语音标记因其存储效率和与大语言模型的兼容性受到关注,可分为声学标记和语义标记,其中语义标记在自动语音识别(ASR)中更具优势。传统上,无监督K-means聚类被用于从语音基础模型(SFMs)中提取语义标记。近期,基于ASR损失训练的有限标量量化(FSQ)等监督方法出现,适用于语音生成。本文系统比较了监督与无监督语义语音标记在儿童语音识别中的表现。结果表明,监督方法不仅优于无监督方法,甚至意外超越连续表示,在超低码率设置下依然表现良好。这些发现凸显了监督语义标记的优势,为离散语音标记化改进提供了新思路。

原文摘要 · Abstract (English)

Discrete speech tokens have gained attention for their storage efficiency and integration with Large Language Models (LLMs). They are commonly categorized into acoustic and semantic tokens, with the latter being more advantageous for Automatic Speech Recognition (ASR). Traditionally, unsupervised K-means clustering has been used to extract semantic speech tokens from Speech Foundation Models (SFMs). Recently, supervised methods, such as finite scalar quantization (FSQ) trained with ASR loss, have emerged for speech generation. Both approaches leverage pre-trained SFMs, benefiting low-resource tasks such as child ASR. This paper systematically compares supervised and unsupervised semantic speech tokens for child ASR. Results show that supervised methods not only outperform unsupervised ones but even unexpectedly surpass continuous representations, and they perform well even in ultra-low bitrate settings. These findings highlight the advantages of supervised semantic tokens and offer insights for improving discrete speech tokenization.

语音识别离散标记儿童语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。