分离音频能量与语义,提升生成模型训练速度和质量。
PoDAR: Power-Disentangled Audio Representation for Generative Modeling

- 通过随机能量增强与隐空间一致性目标,解耦信号能量与语义内容。
- 训练收敛速度提升约2倍,说话人相似度提高0.055,UTMOS提升0.22。
- 能量通道独立后可仅对语义内容施加条件生成,扩大稳定引导范围。
音频潜在扩散模型的性能主要由生成器表达能力及潜在空间建模能力决定。尽管近期研究多关注生成器改进与音频编码器重建保真度提升,本文表明,通过显式因子解耦可显著提升潜在空间建模能力。我们提出PoDAR(Power-Disentangled Audio Representation)框架,利用随机能量增强与潜在一致性目标,将信号能量与不变语义内容解耦。该分解使潜在空间更易建模,不仅加速下游生成模型收敛,还提升最终性能。在Stable Audio 1.0 VAE搭配F5-TTS生成器时,PoDAR实现约2倍收敛加速,同时在LibriSpeech-PC数据集上,说话人相似度提升0.055,UTMOS提升0.22。此外,将能量分离至专用通道后,可仅对能量无关内容应用条件生成(CFG),有效扩展高尺度下的稳定引导范围。
原文摘要 · Abstract (English)
The performance of audio latent diffusion models is primarily governed by generator expressivity and the modelability of the underlying latent space. While recent research has focused primarily on the former, as well as improving the reconstruction fidelity of audio codecs, we demonstrate that latent modelability can be significantly improved through explicit factor disentanglement. We present PoDAR (Power-Disentangled Audio Representation), a framework that utilizes a randomized power augmentation and latent consistency objective to decouple signal power from invariant semantic content. This factorization makes the latent space easier to model, which both accelerates the convergence of downstream generative models and improves final overall performance. When applied to a Stable Audio 1.0 VAE with an F5-TTS generator, PoDAR achieves about a $2\times$ acceleration in convergence to match baseline performance, while increasing final speaker similarity by 0.055 and UTMOS by 0.22 on the LibriSpeech-PC dataset. Furthermore, isolating power into dedicated channels enables the application of CFG exclusively to power-invariant content, effectively extending the stable guidance regime to higher scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。