arXiv:2410.11025eess.AScs.SD2024-10被引 5

提升音频神经编码器的稳定性,让反复压缩不丢质量

Code Drift: Towards Idempotent Neural Audio Codecs

  • 通过微调提升编码器在多次压缩下的输出一致性
  • 三轮编码后,改进模型音频失真显著降低
  • 适合需要反复编辑或压缩的音频生成场景

神经音频编码器在低码率下实现了高质量音频压缩,其基于标记的表示对生成建模很有帮助。尽管压缩比和听觉透明度已有大量研究,但编码器的另一重要特性——恒等性(idempotence)却未受关注:即多次编码后输出是否稳定。我们发现当前先进编码器在多次编码后表现各异,部分模型仅经三轮编码就导致音频质量明显下降。通过分析原因并提出微调方法,我们提升了编码器的恒等性。实验表明,在不损害下游生成性能的前提下,可实现更高恒等性,这对实际文件压缩和迭代生成工作流具有潜在价值。

原文摘要 · Abstract (English)

Neural codecs have demonstrated strong performance in high-fidelity compression of audio signals at low bitrates. The token-based representations produced by these codecs have proven particularly useful for generative modeling. While much research has focused on improvements in compression ratio and perceptual transparency, recent works have largely overlooked another desirable codec property -- idempotence, the stability of compressed outputs under multiple rounds of encoding. We find that state-of-the-art neural codecs exhibit varied degrees of idempotence, with some degrading audio outputs significantly after as few as three encodings. We investigate possible causes of low idempotence and devise a method for improving idempotence through fine-tuning a codec model. We then examine the effect of idempotence on a simple conditional generative modeling task, and find that increased idempotence can be achieved without negatively impacting downstream modeling performance -- potentially extending the usefulness of neural codecs for practical file compression and iterative generative modeling workflows.

音频编码神经编码器恒等性生成建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。