研究发现语音离散表示会丢失声调信息,尤其在低资源语言中。
Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?
- 用k-means对SSL模型隐变量聚类生成离散符号
- 离散符号在普通话和约鲁巴语中显著丢失声调信息
- 建议声调相关任务需考虑任务感知的离散化
语音的离散表示广泛应用于自监督学习(SSL)基础模型,尤其在下游任务数据有限时,如低资源语言场景。通常通过无监督聚类(如k-means)将SSL模型的隐变量转化为符号序列。本研究评估了基于HuBERT base、MandarinHuBERT或XLS-R模型提取的离散符号,能否有效捕捉普通话和约鲁巴语中的声调差异。实验对比了隐向量与离散符号在元音和声调分类任务上的表现,结果表明,即使使用语言特化模型,离散化仍会导致声调信息大幅损失。研究建议离散化过程应具备任务感知性,尤其针对依赖声调的下游任务。
原文摘要 · Abstract (English)
Discrete representations of speech, obtained from Self-Supervised Learning (SSL) foundation models, are widely used, especially where there are limited data for the downstream task, such as for a low-resource language. Typically, discretization of speech into a sequence of symbols is achieved by unsupervised clustering of the latents from an SSL model. Our study evaluates whether discrete symbols - found using k-means - adequately capture tone in two example languages, Mandarin and Yoruba. We compare latent vectors with discrete symbols, obtained from HuBERT base, MandarinHuBERT, or XLS-R, for vowel and tone classification. We find that using discrete symbols leads to a substantial loss of tone information, even for language-specialised SSL models. We suggest that discretization needs to be task-aware, particularly for tone-dependent downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。