arXiv:2409.07858eess.AScs.LG2024-09被引 1

用逆问题求解思路,提升音频解码质量与效率。

Audio Decoding by Inverse Problem Solving

  • 将音频解码建模为逆问题,通过扩散后验采样实现。
  • 在不同码率下均优于传统方法,尤其对钢琴音效提升明显。
  • 降低梯度计算量,适合追求高效高保真音频重建的场景。

本文将音频解码视为逆问题,通过扩散后验采样求解。针对变换域感知音频编码器提供的输入信号测量值,设计了显式条件函数。通过评估一系列码率与任务无关先验模型的任意组合,验证了方法的有效性。例如,用语音-钢琴联合训练模型替代纯语音模型,在显著提升钢琴音质的同时保持了语音表现。使用更通用的音乐模型时,对多种内容类型和码率均获得优于传统方法的解码效果。所提出的噪声均值模型,使扩散后验采样所需的梯度评估次数大幅减少,相比基于Tweedie均值的方法有显著改进。将Tweedie均值与我们的条件函数结合,进一步提升了客观性能。音频演示见https://dpscodec-demo.github.io/。

原文摘要 · Abstract (English)

We consider audio decoding as an inverse problem and solve it through diffusion posterior sampling. Explicit conditioning functions are developed for input signal measurements provided by an example of a transform domain perceptual audio codec. Viability is demonstrated by evaluating arbitrary pairings of a set of bitrates and task-agnostic prior models. For instance, we observe significant improvements on piano while maintaining speech performance when a speech model is replaced by a joint model trained on both speech and piano. With a more general music model, improved decoding compared to legacy methods is obtained for a broad range of content types and bitrates. The noisy mean model, underlying the proposed derivation of conditioning, enables a significant reduction of gradient evaluations for diffusion posterior sampling, compared to methods based on Tweedie's mean. Combining Tweedie's mean with our conditioning functions improves the objective performance. An audio demo is available at https://dpscodec-demo.github.io/.

音频解码扩散模型逆问题高效重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。