用自动生成的文本描述提升声音事件检测准确率
Sound event detection with audio-text models and heterogeneous temporal annotations
- 用自动生成的音频描述作为辅助标签训练模型
- 在强弱标签混合数据下,PSDS-1得分提升至0.218
- 适合处理标注不全、类别极不平衡的声音识别任务
基于音频与相关元数据生成合成字幕的进展,使得自然语言信息可作为其他音频任务的输入。本文提出一种新方法,通过自由形式的文本引导声音事件检测系统。利用机器生成的字幕作为强标签的补充信息进行训练,并评估不同文本输入的效果。此外,研究了仅部分训练数据具有强标签、其余仅具弱时间标签的场景。结果表明,相比传统的CRNN架构,合成字幕在两种情况下均提升了性能:在包含50个高度不平衡类别的数据集上,使用强标签时PSDS-1得分从0.223提升至0.277;当一半数据仅有弱标签时,得分从0.166提升至0.218。
原文摘要 · Abstract (English)
Recent advances in generating synthetic captions based on audio and related metadata allow using the information contained in natural language as input for other audio tasks. In this paper, we propose a novel method to guide a sound event detection system with free-form text. We use machine-generated captions as complementary information to the strong labels for training, and evaluate the systems using different types of textual inputs. In addition, we study a scenario where only part of the training data has strong labels, and the rest of it only has temporally weak labels. Our findings show that synthetic captions improve the performance in both cases compared to the CRNN architecture typically used for sound event detection. On a dataset of 50 highly unbalanced classes, the PSDS-1 score increases from 0.223 to 0.277 when trained with strong labels, and from 0.166 to 0.218 when half of the training data has only weak labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。