构建多标注音频数据集,揭示模型与人耳对声音事件感知的差异
Exploring Differences between Human Perception and Model Inference in Audio Event Recognition
- 通过10名专业标注员构建多标注数据集,量化语义重要性
- 发现人类忽略细微事件,模型易受噪声干扰,检测更敏感
- 为音频事件识别提供贴近人类感知的新评估标准
音频事件识别(AER)传统上关注检测和识别音频事件,但现有模型通常不区分事件在不同上下文中的重要性,导致结果与人类听觉感知存在显著差异。本文引入语义重要性概念,构建多标注前景音频事件识别(MAFAR)数据集,包含由10名专业标注员标注的音频记录。通过标注频率与方差分析,量化语义重要性并研究人类感知机制。对比人类标注与集成预训练模型预测,发现模型在语义识别和事件存在性检测方面均与人类感知存在显著差距:人类倾向于忽略细微或次要事件,而模型易受噪声事件影响;在事件存在性检测中,模型比人类更敏感。
原文摘要 · Abstract (English)
Audio Event Recognition (AER) traditionally focuses on detecting and identifying audio events. Most existing AER models tend to detect all potential events without considering their varying significance across different contexts. This makes the AER results detected by existing models often have a large discrepancy with human auditory perception. Although this is a critical and significant issue, it has not been extensively studied by the Detection and Classification of Sound Scenes and Events (DCASE) community because solving it is time-consuming and labour-intensive. To address this issue, this paper introduces the concept of semantic importance in AER, focusing on exploring the differences between human perception and model inference. This paper constructs a Multi-Annotated Foreground Audio Event Recognition (MAFAR) dataset, which comprises audio recordings labelled by 10 professional annotators. Through labelling frequency and variance, the MAFAR dataset facilitates the quantification of semantic importance and analysis of human perception. By comparing human annotations with the predictions of ensemble pre-trained models, this paper uncovers a significant gap between human perception and model inference in both semantic identification and existence detection of audio events. Experimental results reveal that human perception tends to ignore subtle or trivial events in the event semantic identification, while model inference is easily affected by events with noises. Meanwhile, in event existence detection, models are usually more sensitive than humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。