解析频域依赖方法在声事件检测中的作用机制与效果。
Towards Understanding of Frequency Dependence on Sound Event Detection
- 对比分析滤波增强与频域动态卷积的性能差异。
- 验证频域依赖对声事件检测的关键提升作用。
- 适合关注音频识别与模型可解释性的研究者。
本文深入分析两种频域依赖的声事件检测(SED)方法:FilterAugment 和频域动态卷积(FDY conv)。尽管深度学习技术在其他模式识别领域迅速发展,但其直接迁移至SED常不适用。为此,此前提出两种频域依赖方法:FilterAugment 通过随机加权频带进行数据增强,FDY conv 则采用频适应卷积核的架构。两者在SED中表现优异,本文进一步探究其具体有效性与特性。通过对比不同类别性能,发现两类方法的优劣;利用梯度加权类激活映射(Grad-CAM)分析带与不带频域掩码及两种FilterAugment的模型,揭示其特征关注区域;提出更简化的频域依赖卷积方法并与FDY conv对比,以理解其关键组件;最后通过主成分分析(PCA)展示FDY conv在不同声事件类别上如何动态调整频域卷积核。结果表明,频域依赖在声事件检测中起关键作用,进一步证实了此类方法的有效性。
原文摘要 · Abstract (English)
In this work, we conduct an in-depth analysis of two frequency-dependent methods for sound event detection (SED): FilterAugment and frequency dynamic convolution (FDY conv). The goal is to better understand their characteristics and behaviors in the context of SED. While SED has been rapidly advancing through the adoption of various deep learning techniques from other pattern recognition fields, such adopted techniques are often not suitable for SED. To address this issue, two frequency-dependent SED methods were previously proposed: FilterAugment, a data augmentation randomly weighting frequency bands, and FDY conv, an architecture applying frequency adaptive convolution kernels. These methods have demonstrated superior performance in SED, and we aim to further analyze their detailed effectiveness and characteristics in SED. We compare class-wise performance to find out specific pros and cons of FilterAugment and FDY conv. We apply Gradient-weighted Class Activation Mapping (Grad-CAM), which highlights time-frequency region that is more inferred by the model, on SED models with and without frequency masking and two types of FilterAugment to observe their detailed characteristics. We propose simpler frequency dependent convolution methods and compare them with FDY conv to further understand which components of FDY conv affects SED performance. Lastly, we apply PCA to show how FDY conv adapts dynamic kernel across frequency dimensions on different sound event classes. The results and discussions demonstrate that frequency dependency plays a significant role in sound event detection and further confirms the effectiveness of frequency dependent methods on SED.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。