通过语言查询分离声音事件,提升嘈杂环境下的检测与计数精度
Noise-Robust Sound Event Detection and Counting via Language-Queried Sound Separation
- 基于事件出现检测的计数方法,实现片段与帧级事件统计
- 在高噪声下比现有方法性能更优,尤其在DESED和WildDESED上表现突出
- 适用于需要精准时间定位与抗噪能力的智能音频分析场景
多数声音事件检测(SED)系统在干净数据上表现良好,但在嘈杂环境中性能显著下降。语言查询音频源分离(LASS)模型有望提升鲁棒性,但现有方法需复杂多阶段训练,且缺乏对目标事件的明确引导。为此,我们提出事件出现检测(EAD),一种基于计数的双层次方法,在片段与帧级别统计事件发生次数。基于EAD,我们构建了联合训练的多任务学习框架,同步优化EAD与SED,增强其在噪声环境中的表现。首先,让SED学习与EAD一致的模式;其次,引入任务约束以提高两者预测的一致性。该框架为LASS模型提供更可靠的片段级预测,并强化时间戳检测能力。在DESED与WildDESED数据集上的实验表明,该方法优于现有方法,且在更高噪声水平下优势更明显。
原文摘要 · Abstract (English)
Most sound event detection (SED) systems perform well on clean datasets but degrade significantly in noisy environments. Language-queried audio source separation (LASS) models show promise for robust SED by separating target events; existing methods require elaborate multi-stage training and lack explicit guidance for target events. To address these challenges, we introduce event appearance detection (EAD), a counting-based approach that counts event occurrences at both the clip and frame levels. Based on EAD, we propose a co-training-based multi-task learning framework for EAD and SED to enhance SED's performance in noisy environments. First, SED struggles to learn the same patterns as EAD. Then, a task-based constraint is designed to improve prediction consistency between SED and EAD. This framework provides more reliable clip-level predictions for LASS models and strengthens timestamp detection capability. Experiments on DESED and WildDESED datasets demonstrate better performance compared to existing methods, with advantages becoming more pronounced at higher noise levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。