用数学公式生成声音数据,实现无真实标签的预训练。
Formula-Supervised Sound Event Detection: Pre-Training Without Real Data
- 通过数学公式合成声音信号,以合成参数作为标签进行预训练。
- 在DESED数据集上显著提升模型准确率并加速训练过程。
- 适合缺乏标注数据的声音事件检测研究者使用。
本文提出一种新型公式驱动的监督学习(FDSL)框架,通过公式化方法生成参数化声学信号,用于环境声音分析模型的预训练。针对声音事件检测(SED)任务中真实标注数据稀缺且人工标注存在噪声和主观偏差的问题,我们构建了名为Formula-SED的合成数据集,所有音频均基于数学公式生成。利用每时刻的合成参数作为真实标签,该方法可消除标签噪声与偏倚,支持大规模预训练。实验表明,在DCASE2023 Challenge Task 4所用的DESED数据集上,采用Formula-SED预训练显著提升了模型精度并加快了收敛速度。
原文摘要 · Abstract (English)
In this paper, we propose a novel formula-driven supervised learning (FDSL) framework for pre-training an environmental sound analysis model by leveraging acoustic signals parametrically synthesized through formula-driven methods. Specifically, we outline detailed procedures and evaluate their effectiveness for sound event detection (SED). The SED task, which involves estimating the types and timings of sound events, is particularly challenged by the difficulty of acquiring a sufficient quantity of accurately labeled training data. Moreover, it is well known that manually annotated labels often contain noises and are significantly influenced by the subjective judgment of annotators. To address these challenges, we propose a novel pre-training method that utilizes a synthetic dataset, Formula-SED, where acoustic data are generated solely based on mathematical formulas. The proposed method enables large-scale pre-training by using the synthesis parameters applied at each time step as ground truth labels, thereby eliminating label noise and bias. We demonstrate that large-scale pre-training with Formula-SED significantly enhances model accuracy and accelerates training, as evidenced by our results in the DESED dataset used for DCASE2023 Challenge Task 4. The project page is at https://yutoshibata07.github.io/Formula-SED/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。