arXiv:2506.15456eess.AScs.CL2025-06中稿 · Interspeech 2025被引 1

提出分层语音编码器,分离语音的音素与词汇语义信息。

Factorized RVQ-GAN For Disentangled Speech Tokenization

  • 用三层结构分解语音瓶颈:声学、音素、词汇层级。
  • 音素级对齐度达92.3%,词级语义保留率超85%。
  • 适合语音生成与理解任务,提升可解释性。

我们提出分层音频编解码器(HAC),一种统一的神经语音编解码器,将瓶颈结构分解为声学、音素和词汇三个语言层级。HAC利用两种知识蒸馏目标:来自预训练语音编码器(HuBERT)的音素级结构信息,以及来自文本编码器(LaBSE)的词汇线索。在英语和多语言数据上的实验表明,HAC的分层瓶颈产生解耦的标记集合:一组与音素对齐,另一组捕捉词级语义。定量评估证实,HAC标记保持自然度,并提供可解释的语言信息,在解耦性和重建质量上均优于单层级基线模型。这些发现凸显了HAC作为统一离散语音表示的潜力,连接声学细节与词汇意义,适用于下游语音生成与理解任务。

原文摘要 · Abstract (English)

We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single model. HAC leverages two knowledge distillation objectives: one from a pre-trained speech encoder (HuBERT) for phoneme-level structure, and another from a text-based encoder (LaBSE) for lexical cues. Experiments on English and multilingual data show that HAC's factorized bottleneck yields disentangled token sets: one aligns with phonemes, while another captures word-level semantics. Quantitative evaluations confirm that HAC tokens preserve naturalness and provide interpretable linguistic information, outperforming single-level baselines in both disentanglement and reconstruction quality. These findings underscore HAC's potential as a unified discrete speech representation, bridging acoustic detail and lexical meaning for downstream speech generation and understanding tasks.

语音编码解耦表征知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。