arXiv:2503.12115cs.SDcs.AI2025-03中稿 · IEEE Journal of Se…被引 8

统一语音语义与声纹信息,实现更自然的语音生成

Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations

  • 用低比特神经编解码器学习全局与局部离散表示
  • 生成语音时保留语义和说话人特征,长时一致性好
  • 适合语音合成、跨域语音转换等需要声纹保持的任务

当前大型语音语言模型主要依赖自监督学习表征的语义离散化令牌和神经编解码器生成的声学令牌,遵循语义建模与声学合成范式。然而,语义令牌会丢失说话人特有的副语言特征,而基于提示的声学合成在恢复副语言细节方面存在局限,尤其在提示与目标领域差异较大时鲁棒性差。本文提出UniCodec,一种统一语音令牌学习方法,将语音的全部语义(包括语言与副语言信息)封装进紧凑且语义解耦的统一令牌中。该统一令牌既有助于语音语言模型理解时获取副语言线索,也支持高质量语音生成。通过低比特率神经编解码器,在全局与局部尺度上学习此类解耦离散表示,并融合自监督特征知识。多语言数据集上的大量评估表明,该方法在多个语音处理任务中能生成自然、富有表现力且长期一致的语音输出,同时有效保留副语言属性。

原文摘要 · Abstract (English)

Current large speech language models are mainly based on semantic tokens from discretization of self-supervised learned representations and acoustic tokens from a neural codec, following a semantic-modeling and acoustic-synthesis paradigm. However, semantic tokens discard paralinguistic attributes of speakers that is important for natural spoken communication, while prompt-based acoustic synthesis from semantic tokens has limits in recovering paralinguistic details and suffers from robustness issues, especially when there are domain gaps between the prompt and the target. This paper unifies two types of tokens and proposes the UniCodec, a universal speech token learning that encapsulates all semantics of speech, including linguistic and paralinguistic information, into a compact and semantically-disentangled unified token. Such a unified token can not only benefit speech language models in understanding with paralinguistic hints but also help speech generation with high-quality output. A low-bitrate neural codec is leveraged to learn such disentangled discrete representations at global and local scales, with knowledge distilled from self-supervised learned features. Extensive evaluations on multilingual datasets demonstrate its effectiveness in generating natural, expressive and long-term consistent output quality with paralinguistic attributes well preserved in several speech processing tasks.

语音生成编码器副语言特征统一表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。