arXiv:2510.16841eess.AScs.SD2025-10ACL被引 14

SAC通过双流分离语义与声学特征,提升语音编码质量与可控制性。

SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization

  • 采用语义-声学双流架构,分别优化语义与声学建模
  • 在不同码率下均实现高自然度(UTMOS)和高可懂度(WER)
  • 适合用于大模型语音合成,支持单阶段自回归生成

将连续语音信号转换为离散标记的语音编解码器对语音语言模型至关重要。然而,现有编解码器难以兼顾高质量重建与丰富的语义表示,限制了其在生成与理解任务中的表现。本文提出SAC,一种基于语义-声学双流量化机制的神经语音编解码器。通过将语义与声学建模分离至两个独立流,使每一流可针对其任务进行优化。综合评估表明,SAC在多种码率下均表现出色,尤其在干净与噪声环境下,其UTMOS与WER得分显著领先,体现出优异的自然度与可懂度。此外,SAC在语义表示方面大幅超越已有编解码器,接近连续自监督嵌入水平。将其作为基于大语言模型的文本到语音系统的分词器时,可实现单阶段自回归语音合成模型,性能明显优于当前最优自回归系统。消融分析进一步验证了双流设计的有效性,为可控语音生成提供了新可能。

原文摘要 · Abstract (English)

Speech codecs that convert continuous speech signals into discrete tokens have become essential for speech language models. However, existing codecs struggle to balance high-quality reconstruction with semantically rich representations, limiting their effectiveness in both generative and understanding tasks. In this work, we propose SAC, a neural speech codec with semantic-acoustic dual-stream quantization. By disentangling semantic and acoustic modeling into two dedicated streams, SAC enables each to be optimized for its respective role. Comprehensive evaluations show that SAC achieves strong reconstruction performance across diverse bitrates under both clean and noisy conditions, with particularly high scores on UTMOS and WER, indicating superior naturalness and intelligibility. Moreover, SAC substantially surpasses prior codecs in semantic representation, approaching the level of continuous self-supervised embeddings. When used as a tokenizer for LLM-based text-to-speech, SAC enables a single-stage autoregressive (AR) TTS model that clearly outperforms state-of-the-art AR systems. Our disentanglement analysis further validates the effectiveness of the dual-stream design, offering new potential for controllable speech generation.

语音编码双流结构语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。