arXiv:2606.06740cs.SDcs.AI2026-06中稿 · Interspeech 2026

分析离散语音单元在多语言多说话人生成中的混淆问题

Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations

论文配图:Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations
图 1 · 摘自论文原文
  • 用聚类后离散语音单元构建声码器,研究其在四语种印度语中的表现
  • 小聚类数导致跨语言音素混淆,大聚类提升发音可辨性,显式说话人条件是防身份坍塌关键
  • 适合语音合成、音频大模型研究者,关注多语言语音生成质量优化

通过k-means对自监督嵌入进行聚类获得的离散语音单元会同时携带音位、说话人和语言信息,导致多语言多说话人语音生成中出现说话人混叠与跨语言干扰。尽管这类单元声码器在音频大模型和语音到语音系统中应用日益广泛,但相关研究仍不充分。本文以BigVGAN为基础,对四种印度语言的单元声码器进行了系统分析。通过词错误率(WER)、说话人相似度及单元级指标,研究了聚类规模与条件策略之间的交互作用。结果表明,聚类规模决定了语音可懂性,通过提升音位可区分性实现;显式说话人条件对于防止身份坍塌至关重要。语言监督在低聚类规模下带来额外增益,此时单元仍存在模糊性。较小的单元词表中,不同语言的相似音素会合并到同一簇编号,而较大的聚类逐渐将它们分离。

原文摘要 · Abstract (English)

Discrete speech units obtained via k-means clustering of self supervised embeddings entangle phonetic, speaker, and language information, causing speaker mixing and cross-lingual interference in multilingual multi-speaker speech generation. Despite growing use in Audio LLMs and speech to speech systems, unit vocoders remain underexplored. We analyze a BigVGAN based unit vocoder, across four Indian languages. We study the interaction between cluster size and conditioning strategies using WER, speaker similarity, and unit level metrics. Results show that cluster size governs intelligibility by improving phonetic discriminability, while explicit speaker conditioning is indispensable for preventing identity collapse. Language supervision yields further gains mainly at lower cluster sizes where units remain ambiguous. Our analysis shows similar phonemes across languages collapse to the same cluster IDs at smaller inventories, with larger clusters progressively separating them.

语音合成多语言离散表示声码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。