arXiv:2607.26350cs.SDcs.CL2026-07中稿 · Interspeech2026

发现语音编码器可跨语言复用,但预训练语言需匹配目标语言。

Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens

论文配图:Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens
图 1 · 摘自论文原文
  • 固定编码器训练语言,改变自监督预训练语言测试敏感性。
  • 下游任务性能对编码器语言不敏感,但严重依赖预训练语言。
  • 提示:跨语言应用时只需更换预训练语言,无需重训编码器。

神经音频编码器(NAC)作为离散语音表征的生成工具日益流行。除压缩外,其离散令牌还可用于训练自监督学习(SSL)模型,此类模型称为基于编码器的SSL模型,能显著降低数据存储与计算开销,支持大规模预训练。然而,其对语言变化的敏感性尚不明确。当语言切换时,基于编码器的SSL模型可能需重新训练,削弱了效率优势。本文通过系统分析,固定一个环节的语言(编码器训练语言或SSL预训练语言),变化另一环节,探究语言敏感性。实验表明:下游性能对编码器训练语言不敏感,但对SSL预训练语言高度敏感。结果表明,单一编码器可在多语言间复用,而确保预训练语言与目标语言一致至关重要。

原文摘要 · Abstract (English)

Neural audio codecs (NACs) have become popular for obtaining speech representations as discrete tokens. Beyond compression, discrete tokens can be used to train self-supervised learning (SSL) models. Such models, referred to as codec-based SSL models, reduce data storage and computational cost, enabling scalable SSL pre-training. However, their language sensitivity remains unclear. When the language changes, codec-based SSL models may require retraining, which undermines their efficiency. In this paper, we present a systematic analysis of language sensitivity by varying either the NAC training language or the SSL pre-training language while keeping the other fixed. Experimental results show that downstream performance is insensitive to the NAC training language but strongly dependent on the SSL pre-training language. These findings suggest that a single NAC can be reused across languages, while aligning the SSL pre-training language with the target language is crucial.

语音表示自监督学习跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。