用注意力机制提升声音事件检测对瞬时事件的捕捉能力。
Temporal Attention Pooling for Frequency Dynamic Convolution in Sound Event Detection
- 引入时间注意力池化,动态加权时序特征以增强瞬态响应。
- 在SED任务中提升平均PSDS1达3.02%,最高达0.459。
- 特别适合检测铃声、敲门声等瞬时声音事件,通用性强。
近年来,频域动态卷积(FDY conv)通过自适应频率特征提取显著提升了声音事件检测(SED)性能。然而,其依赖的时序平均池化对所有时间帧一视同仁,难以捕捉如警报声、敲门声和语音爆破音等瞬时事件。为此,本文提出时间注意力池化频域动态卷积(TFD conv),以时间注意力池化(TAP)替代原有平均池化。TAP通过三种互补机制实现:时间注意力池化(TA)突出显著特征,速度注意力池化(VA)捕捉瞬时变化,传统平均池化保持对平稳信号的鲁棒性。消融实验表明,与FDY conv相比,TFD conv仅增加14.8%参数量,平均PSDS1提升3.02%。类间方差分析和Tukey HSD检验进一步证实,该方法显著提升了对瞬时事件的检测性能,最高达到PSDS1=0.456,超越现有最优系统。此外,将TAP与多膨胀FDY conv(MDFD conv)结合,取得最佳效果,PSDS1达0.459,验证了时间注意力与多尺度频域适应的互补优势。结果表明,TFD conv是一种兼具瞬态敏感性和整体鲁棒性的通用框架。
原文摘要 · Abstract (English)
Recent advances in deep learning, particularly frequency dynamic convolution (FDY conv), have significantly improved sound event detection (SED) by enabling frequency-adaptive feature extraction. However, FDY conv relies on temporal average pooling, which treats all temporal frames equally, limiting its ability to capture transient sound events such as alarm bells, door knocks, and speech plosives. To address this limitation, we propose temporal attention pooling frequency dynamic convolution (TFD conv) to replace temporal average pooling with temporal attention pooling (TAP). TAP adaptively weights temporal features through three complementary mechanisms: time attention pooling (TA) for emphasizing salient features, velocity attention pooling (VA) for capturing transient changes, and conventional average pooling for robustness to stationary signals. Ablation studies show that TFD conv improves average PSDS1 by 3.02% over FDY conv with only a 14.8% increase in parameter count. Classwise ANOVA and Tukey HSD analysis further demonstrate that TFD conv significantly enhances detection performance for transient-heavy events, outperforming existing FDY conv models. Notably, TFD conv achieves a maximum PSDS1 score of 0.456, surpassing previous state-of-the-art SED systems. We also explore the compatibility of TAP with other FDY conv variants, including dilated FDY conv (DFD conv), partial FDY conv (PFD conv), and multi-dilated FDY conv (MDFD conv). Among these, the integration of TAP with MDFD conv achieves the best result with a PSDS1 score of 0.459, validating the complementary strengths of temporal attention and multi-scale frequency adaptation. These findings establish TFD conv as a powerful and generalizable framework for enhancing both transient sensitivity and overall feature robustness in SED.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。