系统探索扩散模型在语音编码中的设计,提升低码率下音质表现。
On the Design of Diffusion-based Neural Speech Codecs
- 按条件与输出域对扩散模型语音编码进行分类,构建设计框架。
- 在框架内设计新模型,在客观与主观测试中优于现有基线。
- 为低码率语音编码提供可复现的扩散模型设计路径,适合音频生成研究者。
最近,以生成模型训练的神经语音编码器(NSCs)在低比特率下相比传统编码器展现出更优性能。尽管多数先进NSCs采用生成对抗网络(GAN),但扩散模型(DMs)因其在图像生成中相较GAN的卓越表现,成为有前景的替代方案,并已在音频与语音编码等众多音频生成任务中成功应用。然而,扩散模型用于语音编码的设计尚未系统化。本文通过三项贡献填补该空白:首先,提出基于条件与输出域的分类框架,明确扩散模型语音编码的设计空间并归类已有方法;其次,基于该框架系统探索未被研究的设计,构建并评估新型扩散模型语音编码器;最后,通过客观指标与主观听觉测试对比所提模型与现有GAN及DM基线。结果表明,所提方法在低码率下显著提升音质,验证了框架的有效性与可扩展性。
原文摘要 · Abstract (English)
Recently, neural speech codecs (NSCs) trained as generative models have shown superior performance compared to conventional codecs at low bitrates. Although most state-of-the-art NSCs are trained as Generative Adversarial Networks (GANs), Diffusion Models (DMs), a recent class of generative models, represent a promising alternative due to their superior performance in image generation relative to GANs. Consequently, DMs have been successfully applied for audio and speech coding among various other audio generation applications. However, the design of diffusion-based NSCs has not yet been explored in a systematic way. We address this by providing a comprehensive analysis of diffusion-based NSCs divided into three contributions. First, we propose a categorization based on the conditioning and output domains of the DM. This simple conceptual framework allows us to define a design space for diffusion-based NSCs and to assign a category to existing approaches in the literature. Second, we systematically investigate unexplored designs by creating and evaluating new diffusion-based NSCs within the conceptual framework. Finally, we compare the proposed models to existing GAN and DM baselines through objective metrics and subjective listening tests.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。