arXiv:2502.15430eess.SPcs.SD2025-02被引 1

用谱图最优传输生成音效中间过渡信号,更自然且计算更快。

Audio signal interpolation using optimal transportation of spectrograms

  • 基于谱图的Wasserstein均值全局插值,非逐帧操作
  • 设计新型代价矩阵,禁止时间轴上能量远距离移动
  • 适合音乐合成与环境音效生成,计算效率更高

我们提出一种新方法,用于生成在给定源音和目标音之间插值的合成音频信号。该方法基于源与目标谱图的Wasserstein均值计算,随后进行相位重建与逆变换。与以往方法不同,本方法对谱图进行全局处理,而非逐时帧操作。另一贡献是设计了特定结构的运输代价矩阵,禁止能量在时间轴上远距离转移,并借助非平衡传输框架实现高效最优传输。该代价矩阵在音频层面具有合理性,同时降低计算开销。通过合成乐音与真实环境声音的实验验证了该方法的有效性。

原文摘要 · Abstract (English)

We present a novel approach for generating an artificial audio signal that interpolates between given source and target sounds. Our approach relies on the computation of Wasserstein barycenters of the source and target spectrograms, followed by phase reconstruction and inversion. In contrast with previous works, our new method considers the spectrograms globally and does not operate on a temporal frame-to-frame basis. Another contribution is to endow the transportation cost matrix with a specific structure that prohibits remote displacements of energy along the time axis, and for which optimal transport is made possible by leveraging the unbalanced transport framework. The proposed cost matrix makes sense from the audio perspective and also allows to reduce the computation load. Results with synthetic musical notes and real environmental sounds illustrate the potential of our novel approach.

音频生成最优传输谱图插值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。