让音频编码区分语音、音乐等不同声源,提升生成可控性。
Learning Source Disentanglement in Neural Audio Codec
- 联合学习音频重建与声源分离,显式划分不同声源的代码本。
- 在保持高质量重建的同时,实现潜空间中声源的有效解耦。
- 适合需要精细控制音频生成的场景,如音乐创作或语音合成。
神经音频编码器通过将连续音频信号高效转换为离散标记,显著提升了音频压缩性能,并支持基于这些标记的生成模型进行复杂声音生成。然而,现有神经编码模型通常在大规模、未区分的音频数据集上训练,忽略了语音、音乐和环境音效等声音域之间的本质差异。这一疏忽导致数据建模困难,也增加了声音生成的不可控性。为此,我们提出源解耦神经音频编码器(SD-Codec),该方法结合音频编码与源分离技术,通过联合学习音频重建与分离任务,显式地将不同声源分配至独立的代码本(即离散表示集合)。实验结果表明,SD-Codec不仅保持了竞争性的重建质量,且在分离结果支持下,成功实现了潜空间中不同声源的解耦,从而增强了音频编码的可解释性,并为音频生成提供了更精细的控制潜力。
原文摘要 · Abstract (English)
Neural audio codecs have significantly advanced audio compression by efficiently converting continuous audio signals into discrete tokens. These codecs preserve high-quality sound and enable sophisticated sound generation through generative models trained on these tokens. However, existing neural codec models are typically trained on large, undifferentiated audio datasets, neglecting the essential discrepancies between sound domains like speech, music, and environmental sound effects. This oversight complicates data modeling and poses additional challenges to the controllability of sound generation. To tackle these issues, we introduce the Source-Disentangled Neural Audio Codec (SD-Codec), a novel approach that combines audio coding and source separation. By jointly learning audio resynthesis and separation, SD-Codec explicitly assigns audio signals from different domains to distinct codebooks, sets of discrete representations. Experimental results indicate that SD-Codec not only maintains competitive resynthesis quality but also, supported by the separation results, demonstrates successful disentanglement of different sources in the latent space, thereby enhancing interpretability in audio codec and providing potential finer control over the audio generation process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。