用文本生成音频并精准控制时间,提升声音事件检测效果。
SynSonic: Augmenting Sound Event Detection through Text-to-Audio Diffusion ControlNet and Effective Sample Filtering
- 用文本驱动的扩散模型生成带时间结构的声音事件
- 联合双分类器过滤,确保生成样本高质量
- 适合需要增强数据多样性的声音检测研究者
由于时间标注数据稀缺,数据合成与增强对声音事件检测(SED)至关重要。现有方法如SpecAugment和Mix-up受限于原始样本多样性。近年来的生成模型虽提供新可能,但因缺乏精确时间标注且易引入噪声而难以直接用于SED。为此,我们提出SynSonic,一种专为SED设计的数据增强方法。该方法利用文本到音频的扩散模型,并通过能量包络ControlNet实现时间一致性控制;采用双分类器联合评分策略进行有效样本筛选,并探索其在训练流程中的实际集成。实验表明,SynSonic显著提升了多音符声音检测分数(PSDS1和PSDS2),在时间定位与声类区分上均有改善。
原文摘要 · Abstract (English)
Data synthesis and augmentation are essential for Sound Event Detection (SED) due to the scarcity of temporally labeled data. While augmentation methods like SpecAugment and Mix-up can enhance model performance, they remain constrained by the diversity of existing samples. Recent generative models offer new opportunities, yet their direct application to SED is challenging due to the lack of precise temporal annotations and the risk of introducing noise through unreliable filtering. To address these challenges and enable generative-based augmentation for SED, we propose SynSonic, a data augmentation method tailored for this task. SynSonic leverages text-to-audio diffusion models guided by an energy-envelope ControlNet to generate temporally coherent sound events. A joint score filtering strategy with dual classifiers ensures sample quality, and we explore its practical integration into training pipelines. Experimental results show that SynSonic improves Polyphonic Sound Detection Scores (PSDS1 and PSDS2), enhancing both temporal localization and sound class discrimination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。