arXiv:2604.07467cs.CLcs.LG2026-04中稿 · Speech Prosody 202…

量化语音单元会弱化声调信息,影响中文和约鲁巴语的语音表征质量。

Lexical Tone is Hard to Quantize: Probing Discrete Speech Units in Mandarin and Yorùbá

  • 用多种量化方法从自监督模型中提取离散语音单元
  • 声调在量化后丢失严重,远不如音素结构稳定
  • 建议分步聚类:先分音素,再对残差编码声调

离散语音单元(DSUs)通过量化自监督学习(SSL)模型的表示获得,广泛用于各类语音任务,尤其适用于需联合建模文本与语音的场景。然而我们发现,尽管原始的SSL隐含表示能较好编码声调,但经量化后的DSUs更关注音素结构,导致声调信息编码不稳定。我们在普通话和约鲁巴语两种声调语言上验证了这一现象,且该问题在多种量化方法(包括K-means)下均存在。这表明当前的量化策略对超音段特征(如声调、韵律)存在固有局限。我们提出一种改进方案:先对原始表示进行K-means聚类以保留音素信息,再对残差表示进行第二次聚类,可更有效地编码声调信息,提示未来应发展声调感知的量化方法。

原文摘要 · Abstract (English)

Discrete speech units (DSUs) are derived by quantising representations from models trained using self-supervised learning (SSL). They are a popular representation for a wide variety of spoken language tasks, including those where prosody matters. DSUs are especially convenient for tasks where text and speech are jointly modelled, such as text-to-speech and multimodal dialogue systems. But we have found that DSUs encode suprasegmental information less reliably than segmental structure, which we demonstrate in this work using lexical tone, though this limitation likely extends to other suprasegmental features such as prosody. Our investigations using the tone languages Mandarin and Yorùbá show that the SSL latent representations themselves do encode tone, yet DSUs obtained using quantisation tend to prioritise phonetic structure, which makes lexical tone less reliably encoded. This remains true for a variety of quantisation methods, not only the most common, K-means. We conclude that current DSU quantisation strategies have limitations for suprasegmental features, which suggests a need for new, tone-aware (or prosody-aware) techniques in speech representation learning. We point towards a potential form of the solution by performing K-means clustering once to encode phonetic information, then again on the residual representation, which better encodes lexical tone.

语音表征声调量化自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。