用帧级预测生成伪强标签,提升弱监督声音事件检测的定位精度。
Pseudo Strong Labels from Frame-Level Predictions for Weakly Supervised Sound Event Detection
- 从帧级预测生成伪强标签,弥补弱监督数据的时间信息缺失。
- 在三个基准数据集上,PSDS1指标提升最高达7.6%。
- 适合研究弱监督音频识别与时间定位的学者参考。
弱监督声音事件检测(WSSED)依赖仅含音频标签而无精确起止时间的数据,因强标注数据稀缺而广泛使用。本文提出帧级伪强标签生成(FPSL),通过帧级预测生成伪强标签,增强训练时的时间定位能力,解决片段级弱监督的局限性。在DCASE2017 Task 4、DCASE2018 Task 4和UrbanSED三个基准数据集上验证,显著提升多项关键指标:例如,使用卷积循环神经网络(CRNN)的模型在DCASE2017上PSDS1提升4.9%,在DCASE2018上提升7.6%,在UrbanSED上提升1.8%,证实该方法有效提升模型性能。
原文摘要 · Abstract (English)
Weakly Supervised Sound Event Detection (WSSED), which relies on audio tags without precise onset and offset times, has become prevalent due to the scarcity of strongly labeled data that includes exact temporal boundaries for events. This study introduces Frame-level Pseudo Strong Labeling (FPSL) to overcome the lack of temporal information in WSSED by generating pseudo strong labels from frame-level predictions. This enhances temporal localization during training and addresses the limitations of clip-wise weak supervision. We validate our approach across three benchmark datasets (DCASE2017 Task 4, DCASE2018 Task 4, and UrbanSED) and demonstrate significant improvements in key metrics such as the Polyphonic Sound Detection Scores (PSDS), event-based F1 scores, and intersection-based F1 scores. For example, Convolutional Recurrent Neural Networks (CRNNs) trained with FPSL outperform baseline models by 4.9% in PSDS1 on DCASE2017, 7.6% on DCASE2018, and 1.8% on UrbanSED, confirming the effectiveness of our method in enhancing model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。