arXiv:2409.09162eess.AScs.SD2024-09中稿 · ICASSP 2025被引 8

用Mamba模型生成影视音效,速度更快效果更优。

MambaFoley: Foley Sound Generation using Selective State-Space Models

  • 基于扩散模型与Mamba SSM构建音效生成新方法
  • 在客观与主观评测中均优于现有先进模型
  • 适合需要高效音效生成的多媒体制作场景

深度学习的进步推动了音频内容生成技术的广泛应用,尤其在去噪扩散概率模型(DDPM)方面表现突出。其中,音效合成(Foley Sound Synthesis)因其在多媒体内容创作中的关键作用而备受关注。由于声音具有强时间依赖性,设计能有效处理音频序列建模的生成模型至关重要。近年来,选择性状态空间模型(SSMs)被提出作为替代方案,展现出与传统方法相当的性能,同时计算复杂度更低。本文首次将最新提出的Mamba SSM引入音效生成任务,提出MambaFoley——一种基于扩散模型的音效生成方法。通过客观与主观评估,我们验证了该方法的有效性,并与当前最先进的音效生成模型进行对比。

原文摘要 · Abstract (English)

Recent advancements in deep learning have led to widespread use of techniques for audio content generation, notably employing Denoising Diffusion Probabilistic Models (DDPM) across various tasks. Among these, Foley Sound Synthesis is of particular interest for its role in applications for the creation of multimedia content. Given the temporal-dependent nature of sound, it is crucial to design generative models that can effectively handle the sequential modeling of audio samples. Selective State Space Models (SSMs) have recently been proposed as a valid alternative to previously proposed techniques, demonstrating competitive performance with lower computational complexity. In this paper, we introduce MambaFoley, a diffusion-based model that, to the best of our knowledge, is the first to leverage the recently proposed SSM known as Mamba for the Foley sound generation task. To evaluate the effectiveness of the proposed method, we compare it with a state-of-the-art Foley sound generative model using both objective and subjective analyses.

音效生成Mamba扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。