提出分层解耦框架,分离语音语义与声学信息,提升重建质量。
HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement
- 将自监督表示分解为语义与声学残差两层编码,分层处理
- 相比SpeechTokenizer,WER降低44%,比特率减半,重建更优
- 适合需要高质量语音重建和高精度识别的场景
语音建模中的有效表征需在语义相关性与声学保真度间取得平衡,但现有方法难以兼顾。为此,我们提出分层声学与语义表征解耦框架HASRD(发音同'hazard'),将自监督学习表征分解为离散的语义与声学令牌。语义表示由第一代码本承载,声学残差则由后续代码本编码。该设计在保持自动语音识别(ASR)性能的同时,实现高质量语音重建。此外,通过优化编码器效率,在不牺牲重建质量的前提下进一步提升ASR表现。相较于SpeechTokenizer,HASRD实现44%相对WER降低、重建质量更优且比特率降低2倍,充分验证了其在分离声学与语义信息上的有效性。
原文摘要 · Abstract (English)
Effective speech representations for spoken language models must balance semantic relevance with acoustic fidelity for high-quality reconstruction. However, existing approaches struggle to achieve both simultaneously. To address this, we introduce Hierarchical Acoustic and Semantic Representation Disentanglement (HASRD, pronounced `hazard'), a framework that factorizes self-supervised learning representations into discrete semantic and acoustic tokens. HASRD assigns the semantic representation to the first codebook, while encoding acoustic residuals in subsequent codebooks. This preserves ASR performance while achieving high-quality reconstruction. Additionally, we enhance HASRD's encoder efficiency, improving ASR performance without compromising reconstruction quality. Compared to SpeechTokenizer, HASRD achieves a 44% relative WER improvement, superior reconstruction quality, and 2x lower bitrate, demonstrating its effectiveness in disentangling acoustic and semantic information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。