arXiv:2409.08425eess.AScs.SD2024-09被引 19

用语言引导的扩散模型实现精准目标声音提取

SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer

论文配图:SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
图 1 · 摘自论文原文
  • 用带跳跃连接的Transformer替代U-Net,处理音频潜在特征
  • 在FSD Kaggle和AudioSet上达到顶尖性能,零样本/少样本表现强
  • 结合文本到音频生成数据训练,泛化能力出色,适合音频分离场景

本文提出SoloAudio,一种基于扩散模型的目标声音提取(TSE)新方法。该方法在音频潜空间上训练扩散模型,将传统U-Net替换为跳连结构的Transformer,以处理潜变量特征。SoloAudio通过CLAP模型提取目标声音特征,支持音频与语言双模态输入。同时,利用先进文本到音频模型生成的合成音频进行训练,显著提升对域外数据和未见声音事件的泛化能力。在FSD Kaggle 2018混合数据集和AudioSet真实数据上,该方法在域内与域外测试中均达到当前最优性能,展现出优异的零样本与少样本能力。代码与演示已公开。

原文摘要 · Abstract (English)

In this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-connected Transformer that operates on latent features. SoloAudio supports both audio-oriented and language-oriented TSE by utilizing a CLAP model as the feature extractor for target sounds. Furthermore, SoloAudio leverages synthetic audio generated by state-of-the-art text-to-audio models for training, demonstrating strong generalization to out-of-domain data and unseen sound events. We evaluate this approach on the FSD Kaggle 2018 mixture dataset and real data from AudioSet, where SoloAudio achieves the state-of-the-art results on both in-domain and out-of-domain data, and exhibits impressive zero-shot and few-shot capabilities. Source code and demos are released.

声音分离扩散模型语言引导音频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。