用数据增强分离语音语义、音色和韵律,实现低比特率高质量编码。
AugCodec: A Low-Bitrate Disentangled Neural Speech Codec via Data Augmentation

- 通过定制数据增强生成语音变体,分离出语义、音色和韵律三类离散码元。
- 在LibriSpeech上以12.5Hz码率实现更优重建质量与解耦性能。
- 适合低带宽语音传输与个性化语音合成场景。
我们提出AugCodec,一种基于数据增强的低比特率解耦神经语音编解码器,可将语音分解为语义、说话人和韵律三类独立码元。具体地,采用针对性的数据增强策略生成语音变体,分别作为输入提取保留目标属性并抑制其他属性的码元。该解耦策略显著降低码元速率。此外,引入增强损失,对齐源语音与音色转换后语音的语义编码输出,促使生成与说话人无关的嵌入表示,缓解音色转换带来的声学失配问题。在LibriSpeech test-clean上的实验表明,AugCodec在重建质量和解耦性能上均显著优于现有最先进方法,且仅需12.5Hz的码元速率。
原文摘要 · Abstract (English)
We propose AugCodec, a low-bitrate disentangled neural speech codec that leverages data augmentation to decompose speech into three distinct components: semantic, speaker, and prosody tokens. Specifically, we employ tailored augmenta tion strategies to transform speech into distinct variants, each serving as input for extracting tokens that preserve the target attribute while suppressing others. This disentanglement strategy enables substantial reduction in token rate. Further more, we introduce an augmentation loss that aligns semantic encoder outputs between source and voice-converted speech, encouraging speaker-agnostic embeddings while mitigating the acoustic mismatch induced by voice conversion. Experiments on LibriSpeech test-clean demonstrate that AugCodec significantly outperforms state-of-the-art methods in both reconstruction quality and disentanglement, while operating at only 12.5Hz with three token streams.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。