arXiv:2607.29148eess.AS2026-07

用双路径结构高效生成逼真音效,小模型媲美大模型

Exploring Efficient Waveform Diffusion Models for Foley Sound Generation

论文配图:Exploring Efficient Waveform Diffusion Models for Foley Sound Generation
图 1 · 摘自论文原文
  • 在时频域并行做子带与帧级自注意力,兼顾精度与效率
  • 300万参数模型性能媲美5000万参数模型
  • 适合资源受限场景下的音效生成应用

近年来,扩散模型已在波形空间实现高保真音效生成。现有波形扩散模型多采用时域架构(如基于CNN的U-Net和DiffWave)或频域Transformer建模时序依赖,但普遍模型庞大、计算成本高,轻量高效架构仍待探索。本文提出一种双路径(DP)波形扩散架构,在时频域沿子带和帧轴分别执行维度自注意力,实现精细的时空谱建模同时保持高效。基于此设计,我们开发了两种变体:DP-DiT与DP-U-Net。在DCASE与FSD-Kaggle2018数据集上的实验表明,其性能优异;尤其300万参数版本表现可比超过5000万参数的模型。音频样例见https://samplesdemo.github.io/DP-Foley/。

原文摘要 · Abstract (English)

Recent advances in diffusion models have enabled high-fidelity Foley sound generation directly in the waveform space. Existing waveform diffusion models primarily rely on time-domain architectures, such as CNN-based U-Nets and DiffWave-style models, or frequency-domain Transformers modeling temporal dependencies. However, these systems are typically built with large model capacities and substantial computational costs, leaving compact and efficient waveform diffusion architectures largely underexplored. In this work, we introduce a Dual-Path (DP) architecture for waveform diffusion that performs dimension-wise self-attention along both subband and frame axes in the time-frequency domain. This DP design enables fine-grained temporal-spectral modeling while maintaining high efficiency. Based on the proposed DP backbone, we develop two variants: DP-DiT and DP-U-Net. Experiments on the DCASE and FSD-Kaggle2018 datasets demonstrate their superior performance. Notably, the 3M parameter variant achieves performance comparable to models with more than 50M parameters. Audio samples are available at https://samplesdemo.github.io/DP-Foley/.

音效生成扩散模型轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。