DS-Codec通过双阶段架构切换,实现更高质量的语音重建。
DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec
- 采用镜像与非镜像结构交替训练,提升编码器鲁棒性。
- 实验表明该策略显著改善高保真语音重建效果。
- 适合语音合成与语音压缩领域的研究者参考。
神经语音编解码器对推动文本到语音(TTS)系统发展至关重要。随着大语言模型在文本生成中的成功,开发高质量的语音分词器变得日益重要。本文提出DS-Codec,一种新型神经语音编解码器,采用双阶段训练框架,结合镜像与非镜像架构的动态切换,旨在实现更优的语音重建。我们进行了大量实验与消融研究,评估该训练策略的有效性,并对比两种架构的性能表现。结果表明,镜像结构显著增强了学习到的码本鲁棒性,而训练策略则平衡了镜像与非镜像结构的优势,从而实现了更高质量的高保真语音重建。
原文摘要 · Abstract (English)
Neural speech codecs are essential for advancing text-to-speech (TTS) systems. With the recent success of large language models in text generation, developing high-quality speech tokenizers has become increasingly important. This paper introduces DS-Codec, a novel neural speech codec featuring a dual-stage training framework with mirror and non-mirror architectures switching, designed to achieve superior speech reconstruction. We conduct extensive experiments and ablation studies to evaluate the effectiveness of our training strategy and compare the performance of the two architectures. Our results show that the mirrored structure significantly enhances the robustness of the learned codebooks, and the training strategy balances the advantages between mirrored and non-mirrored structures, leading to improved high-fidelity speech reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。