用简单向量相加就能实现音乐带宽扩展,效果媲美复杂模型。
On the Geometry of Music Bandwidth Extension in Latent Spaces of Audio Codecs

- 在音频编码器潜空间中,通过计算干净与受损样本中心的迁移向量并相加,实现修复。
- 该方法在多个神经编码器上表现接近扩散模型,峰值信噪比提升1.2~2.3dB。
- 适合关注模型效率、潜空间结构或想建立基准的音频修复研究者。
当前音频修复越来越多依赖大规模条件潜空间生成模型(如扩散模型、薛定谔桥、流匹配),用于恢复带宽限制或噪声等退化问题。本文分析了多种前沿方法在多个神经音频编码器潜空间中对音乐带宽扩展的表现,发现仅需在参考集上估计干净与受损潜空间中心之间的单个传输向量,并将其加至受损潜空间,即可达到与大型扩散模型相当的修复效果。这表明:第一,部分神经编码器潜空间具有与音频带宽对齐的结构;第二,在此类情况下,复杂条件模型带来的增益有限。我们据此提出,未来研究可利用潜空间结构以提升训练与参数效率,并改善整体性能。此外,建议将此简单算术变换作为音乐带宽扩展研究的基线,以评估可学习参数对修复效果的实际贡献。
原文摘要 · Abstract (English)
Recent audio restoration increasingly relies on large-scale conditional latent generative modeling, including diffusion, Schrodinger Bridges, and Flow Matching variants, to invert degradations such as bandwidth limitation or noise. We present an analysis of the performance of various state-of-the-art methods compared to simple arithmetic transformations in the latent spaces of multiple neural codecs for musical bandwidth extension. We show that estimating a single transport vector between the clean and degraded latent centroids on a reference set, and adding it to degraded latents, can yield restoration performance competitive with large diffusion models. This suggests, first, that some neural codec latent spaces exhibit structure aligned with audio bandwidth; and second, that in such cases complex conditional models may offer only limited gains over a simple vector addition. We argue that these findings reveal an interesting avenue for future research whereby models could take advantage of the latent space structure in order to offer greater training and parameter efficiency, and overall better performance. Additionally, we propose to consider this simple arithmetic transformation as a baseline for music bandwidth extension research, as it allows an assessment of the contribution of learnable parameters towards restoration performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。