用混合方法提升音频事件检测,仅靠少量标注数据就达到新纪录。
Semi-Supervised Sound Event Detection with Conditional Mixup and Embedding-Level Contrastive Loss

- 设计条件混合策略,同时利用数据组合与扰动增强模型鲁棒性。
- 在DESED验证集上达0.645 PSDS1和0.822 PSDS2,刷新最佳性能。
- 适合资源有限但有大量未标注音频的场景,如智能监控、环境监测。
声音事件检测(SED)是声学环境分析的核心模块,但其性能常受限于标注数据稀缺。现有系统依赖大规模预训练音频基础模型,然而有效微调仍具挑战,因标注数据少而无标注数据丰富。先前工作ATST-SED采用伪标签驱动的半监督微调框架。本文进一步改进该框架,引入受ATST-Frame预训练启发的嵌入级自监督对比损失,更充分挖掘微调阶段的无标注数据。一个关键挑战是:混合方法在两个目标中角色不同——伪标签学习使用组合型混合,而对比学习将其视为扰动。为解决此矛盾,我们提出条件混合策略,将组合型与扰动型混合统一于同一半监督框架,并定义相应的嵌入级对比损失。所提模型在DESED验证集上取得0.645 PSDS1与0.822 PSDS2,建立新基准。
原文摘要 · Abstract (English)
Sound event detection (SED) is a core module for acoustic environmental analysis, yet its performance is often limited by scarce labeled data. Recent systems leverage large pretrained audio foundation models, but effective fine-tuning remains challenging because labeled data are limited while unlabeled data are abundant. A previous work, ATST-SED, addressed this problem with a pseudo-label based semi-supervised fine-tuning framework. In this work, we further improve the framework by adopting an embedding-level self-supervised contrastive loss inspired by ATST-Frame pretraining. This contrastive objective better exploits unlabeled data during fine-tuning. One challenge is that mixup serves different roles in the two objectives: pseudo-label learning uses composition mixup, while contrastive learning treats mixup as a perturbation. To resolve this mismatch, we propose conditional mixup, which combines composition mixup and perturbation mixup in one semi-supervised framework and defines the corresponding embedding-level contrastive losses. The resulting model achieves 0.645 PSDS1 and 0.822 PSDS2 on the DESED validation set, establishing a new state of the art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。