用复数向量量化实现高效高保真音频编码,无需对抗训练。
EuleroDec: A Complex-Valued RVQ-VAE for Efficient and Robust Audio Coding
- 采用复数域的向量量化自编码器,全程保持幅度与相位耦合。
- 训练步数减少一个数量级,性能仍达最先进水平。
- 适合需要低延迟、高稳定性的音频压缩应用场景。
音频编解码器通过将原始波形压缩至低带宽格式,支撑音乐生成、流媒体和沉浸式媒体应用。近期研究趋向于在频谱域处理,但传统频谱表示难以建模相位信息,该信息天然具有复数特性。现有频域神经编解码器通常忽略相位或将其拆分为两个实值通道,导致空间保真度下降,需依赖对抗判别器或扩散后处理来弥补表征能力不足,影响收敛速度与训练稳定性。本文提出一种端到端的复数域向量量化变分自编码器(EuleroDec),在分析-量化-合成全链路中完整保留幅度与相位耦合关系,无需对抗判别器或扩散后处理。在不使用GAN或扩散模型的前提下,本方法在域内性能达到甚至超越长期训练的基线,在域外任务上也达到最先进水平。相比传统需训练数十万步的模型,本方法将训练预算降低一个数量级,显著提升计算效率的同时保持高感知质量。
原文摘要 · Abstract (English)
Audio codecs power discrete music generative modelling, music streaming and immersive media by shrinking PCM audio to bandwidth-friendly bit-rates. Recent works have gravitated towards processing in the spectral domain; however, spectrogram-domains typically struggle with phase modeling which is naturally complex-valued. Most frequency-domain neural codecs either disregard phase information or encode it as two separate real-valued channels, limiting spatial fidelity. This entails the need to introduce adversarial discriminators at the expense of convergence speed and training stability to compensate for the inadequate representation power of the audio signal. In this work we introduce an end-to-end complex-valued RVQ-VAE audio codec that preserves magnitude-phase coupling across the entire analysis-quantization-synthesis pipeline and removes adversarial discriminators and diffusion post-filters. Without GANs or diffusion we match or surpass much longer-trained baselines in-domain and reach SOTA out-of-domain performance. Compared to standard baselines that train for hundreds of thousands of steps, our model reducing training budget by an order of magnitude is markedly more compute-efficient while preserving high perceptual quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。