用扩散模型实现0.096kbps超低码率音效压缩,重建效果逼真。
S-PRESSO: Ultra Low Bitrate Sound Effect Compression With Diffusion Autoencoders And Offline Quantization
- 基于预训练扩散模型的离线量化编码器,生成连续与离散嵌入。
- 在0.096kbps下实现750倍压缩率,帧率低至1Hz仍保持高保真。
- 优于现有连续与离散方法,在音频质量与相似性上表现更优。
神经音频压缩模型近期实现了极高的压缩率,支持高效的潜在生成建模。相反,潜在生成模型也被用于压缩,推动了连续与离散方法的极限。然而,现有方法仍受限于低分辨率音频,在极低码率下失真明显。本文提出S-PRESSO,一种48kHz音效压缩模型,通过离线量化在超低码率(最低0.096 kbps)下生成连续与离散嵌入。模型依赖预训练潜在扩散模型解码由潜在编码器学习的压缩音频嵌入。利用扩散解码器的生成先验,实现极低帧率(最低1Hz,750倍压缩率),在牺牲精确保真度的前提下生成逼真且自然的重构结果。尽管压缩率极高,实验表明S-PRESSO在音频质量、声学相似性和重建指标上均优于连续与离散基线方法。
原文摘要 · Abstract (English)
Neural audio compression models have recently achieved extreme compression rates, enabling efficient latent generative modeling. Conversely, latent generative models have been applied to compression, pushing the limits of continuous and discrete approaches. However, existing methods remain constrained to low-resolution audio and degrade substantially at very low bitrates, where audible artifacts are prominent. In this paper, we present S-PRESSO, a 48kHz sound effect compression model that produces both continuous and discrete embeddings at ultra-low bitrates, down to 0.096 kbps, via offline quantization. Our model relies on a pretrained latent diffusion model to decode compressed audio embeddings learned by a latent encoder. Leveraging the generative priors of the diffusion decoder, we achieve extremely low frame rates, down to 1Hz (750x compression rate), producing convincing and realistic reconstructions at the cost of exact fidelity. Despite operating at high compression rates, we demonstrate that S-PRESSO outperforms both continuous and discrete baselines in audio quality, acoustic similarity and reconstruction metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。