揭示声音事件检测标注者间标签差异,助力构建鲁棒模型
LEAD Dataset: How Can Labels for Sound Event Detection Vary Depending on Annotators?
- 20名标注者对同一音频片段进行标注,对比标签差异
- 发现同类声音易混淆,起止时间难区分导致标注不一致
- 适合研究标注偏差与提升模型鲁棒性的团队
本文提出大型标注者标签数据集LEAD,用于深入理解声音事件检测(SED)中强标签的标注差异。在实际标注中,由于人工标注耗时且常由多人协作完成,不同标注者对相同音频片段的事件类型和起止时间存在显著差异。尤其当声音类别相似或时间边界模糊时,难以达成一致标签。若训练与测试数据标注者不同,可能导致模型评估失真。LEAD数据集包含来自TUT Sound Events 2016/2017、TUT Acoustic Scenes 2016和URBAN-SED的音频片段,每个片段由20位不同标注者独立标注强标签,可量化分析标注者间差异,为开发对标注变化鲁棒的SED模型提供依据。
原文摘要 · Abstract (English)
In this paper, we introduce a LargE-scale Annotator's labels for sound event Detection (LEAD) dataset, which is the dataset used to gain a better understanding of the variation in strong labels in sound event detection (SED). In SED, it is very time-consuming to collect large-scale strong labels, and in most cases, multiple workers divide up the annotations to create a single dataset. In general, strong labels created by multiple annotators have large variations in the type of sound events and temporal onset/offset. Through the annotations of multiple workers, uniquely determining the strong label is quite difficult because the dataset contains sounds that can be mistaken for similar classes and sounds whose temporal onset/offset is difficult to distinguish. If the strong labels of SED vary greatly depending on the annotator, the SED model trained on a dataset created by multiple annotators will be biased. Moreover, if annotators differ between training and evaluation data, there is a risk that the model cannot be evaluated correctly. To investigate the variation in strong labels, we release the LEAD dataset, which provides distinct strong labels for each clip annotated by 20 different annotators. The LEAD dataset allows us to investigate how strong labels vary from annotator to annotator and consider SED models that are robust to the variation of strong labels. The LEAD dataset consists of strong labels assigned to sound clips from TUT Sound Events 2016/2017, TUT Acoustic Scenes 2016, and URBAN-SED. We also analyze variations in the strong labels in the LEAD dataset and provide insights into the variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。