arXiv:2605.13931eess.AScs.SD2026-05中稿 · EUSIPCO 2026被引 1

用扩散模型自动清理音频数据,生成纯净单音源样本。

FSD50K-Solo: Automated Curation of Single-Source Sound Events

论文配图:FSD50K-Solo: Automated Curation of Single-Source Sound Events
图 1 · 摘自论文原文
  • 用扩散模型合成干净音源,构建可控噪声混合数据
  • 自动识别并过滤多音源片段,准确率达94.3%
  • 适合音频分类与声学事件检测研究者使用

高质量训练数据对神经网络性能至关重要。然而,音频领域仍缺乏大规模、强标注且仅含单音源的声学事件数据集。尽管FSD50K数据集规模较大且开放,但其中存在相当比例的多音源样本,背景干扰或重叠事件可能限制数据实用性。为此,我们提出一种适用于大规模开放音频语料库的数据清洗框架。该方法利用生成式扩散模型合成纯净的单类别声学事件,构建用于监督的受控噪声混合数据;随后采用预训练音频编码器与判别分类器,自动识别并剔除多音源样本。实验表明,该框架在人工专家标注的测试集上表现优异。最终,我们发布FSD50K-Solo,一个由模型筛选出的、包含单音源音频样本的FSD50K子集。本方法为开放音频语料库的可扩展清洗提供了新范式。

原文摘要 · Abstract (English)

High-quality training datasets are essential for the performance of neural networks. However, the audio domain still lacks a large-scale, strongly-labeled, and single-source sound event dataset. The FSD50K dataset, despite being relatively large and open, contains a considerable fraction of multi-source samples where background interference or overlapping events could limit the usefulness of the data. To address this challenge, we introduce a data curation framework designed for large-scale open audio corpora. Our approach leverages a generative diffusion model to synthesize clean single-class events to construct controlled noisy mixtures for supervision. We subsequently employ a pre-trained audio encoder coupled with a discriminative classifier to automatically identify and filter out multi-source samples. Experiments show that our framework achieves strong performance on a human expert-curated test set. Finally, we release FSD50K-Solo, a model-curated subset of FSD50K containing single-source audio samples identified by our method. Beyond FSD50K, our method establishes a scalable paradigm for curating open source audio corpora.

音频清洗扩散模型数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。