提出DC-Spin,让语音模型忽略说话人差异,更好理解与生成语音。
DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
- 用双码本设计,提取不依赖说话人的语音特征。
- 在零样本任务中表现优异,支持流式处理且无需重训练。
- 适合语音识别、语音合成等需要泛化能力的研究者。
语音语言模型(SLMs)随着文本型解码器仅模型的发展日益受到关注。SLMs 同时处理文本与语音,实现语音理解和生成的同步。本文提出双码本说话人无关聚类(DC-Spin),旨在通过连接音频信号与SLM标记来改进语音分词。DC-Spin 提取富含音素信息且对输入变化具有鲁棒性的说话人无关标记,提升零样本SLM任务性能和语音重建质量。我们采用分块策略实现无需重训练的流式处理,避免性能下降。对比自监督方法与神经音频编解码器、模型可扩展性及下游任务代理的结果表明,易于由n-gram语言模型建模或与音素对齐的标记表现更优,为SLM语音分词器的设计提供了重要启示。
原文摘要 · Abstract (English)
Spoken language models (SLMs) have gained increasing attention with advancements in text-based, decoder-only language models. SLMs process text and speech, enabling simultaneous speech understanding and generation. This paper presents Double-Codebook Speaker-invariant Clustering (DC-Spin), which aims to improve speech tokenization by bridging audio signals and SLM tokens. DC-Spin extracts speaker-invariant tokens rich in phonetic information and resilient to input variations, enhancing zero-shot SLM tasks and speech resynthesis. We propose a chunk-wise approach to enable streamable DC-Spin without retraining and degradation. Comparisons of tokenization methods (self-supervised and neural audio codecs), model scalability, and downstream task proxies show that tokens easily modeled by an n-gram LM or aligned with phonemes offer strong performance, providing insights for designing speech tokenizers for SLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。