通过精准建模事件起止点,提升声音事件检测精度与效率
Sound Event Detection with Boundary-Aware Optimization and Inference
- 显式建模事件起始与终止时刻,优化时序定位
- 在AudioSet强标注集上达成新最佳性能,无需后处理调参
- 适合需要高精度时序定位的音频分析任务
时序检测问题广泛存在于时间序列估计、行为识别和声音事件检测(SED)等领域。本文提出一种新型时序事件建模方法,通过显式建模事件起始与终止,并引入边界感知的优化与推理策略,显著提升时序事件检测能力。所提方法包含新的时序建模模块——循环事件检测(RED)和事件提议网络(EPN),结合定制损失函数,实现更高效、更精确的时序事件检测。我们在AudioSet的时序强标注子集上评估该方法,实验结果表明,该方法不仅优于传统帧级SED模型搭配先进后处理的效果,还消除了对后处理超参数调优的需求,并在所有AudioSet Strong类别上达到新的最佳性能。
原文摘要 · Abstract (English)
Temporal detection problems appear in many fields including time-series estimation, activity recognition and sound event detection (SED). In this work, we propose a new approach to temporal event modeling by explicitly modeling event onsets and offsets, and by introducing boundary-aware optimization and inference strategies that substantially enhance temporal event detection. The presented methodology incorporates new temporal modeling layers - Recurrent Event Detection (RED) and Event Proposal Network (EPN) - which, together with tailored loss functions, enable more effective and precise temporal event detection. We evaluate the proposed method in the SED domain using a subset of the temporally-strongly annotated portion of AudioSet. Experimental results show that our approach not only outperforms traditional frame-wise SED models with state-of-the-art post-processing, but also removes the need for post-processing hyperparameter tuning, and scales to achieve new state-of-the-art performance across all AudioSet Strong classes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。