用语义潜变量替代声学潜变量,提升音频生成质量与理解统一性
SemanticVocoder: Bridging Audio Generation and Audio Understanding via Semantic Latents
- 用语义编码器潜变量替代传统声学潜变量进行音频生成
- 在AudioCaps测试集上达到12.823的Frechet距离和1.709的音频弗雷切特距离
- 为音频生成与理解统一于同一语义空间提供新路径,适合语音生成研究者
近期音频生成模型通常依赖变分自编码器(VAEs),在VAE潜空间中进行生成。尽管VAEs在压缩与重建方面表现优异,但其潜变量本质上编码的是低层次声学细节而非语义区分信息,导致事件语义纠缠,增加了生成模型训练难度。为此,我们摒弃了VAE声学潜变量,引入语义编码器潜变量,提出一种直接从语义潜变量合成波形的生成声码器SemanticVocoder。使用SemanticVocoder,我们的文本到音频生成模型在AudioCaps测试集上达到12.823的弗雷切特距离与1.709的弗雷切特音频距离。相较于声学VAE潜变量,所引入的语义潜变量表现出更优的区分能力。该方法不仅提升了生成性能,也为在共享语义空间中统一音频理解与生成提供了有前景的尝试。生成样本见https://zeyuxie29.github.io/SemanticVocoder/。
原文摘要 · Abstract (English)
Recent audio generation models typically rely on Variational Autoencoders (VAEs) and perform generation within the VAE latent space. Although VAEs excel at compression and reconstruction, their latents inherently encode low-level acoustic details rather than semantically discriminative information, leading to entangled event semantics and complicating the training of generative models. To address these issues, we discard VAE acoustic latents and introduce semantic encoder latents, thereby proposing SemanticVocoder, a generative vocoder that directly synthesizes waveforms from semantic latents. Equipped with SemanticVocoder, our text-to-audio generation model achieves a Frechet Distance of 12.823 and a Frechet Audio Distance of 1.709 on the AudioCaps test set, as the introduced semantic latents exhibit superior discriminability compared to acoustic VAE latents. Beyond improved generation performance, it also serves as a promising attempt towards unifying audio understanding and generation within a shared semantic space. Generated samples are available at https://zeyuxie29.github.io/SemanticVocoder/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。