用积分梯度分析声音分类器,发现其能有效定位音频事件的时间边界。
Evaluating the Temporal Detection Capability of Integrated Gradients Applied on Sound Classifier

- 在无时间标注的分类器上应用积分梯度,检测声音事件的起止时间。
- 平均交并比达0.39,帧级F1为0.52,点位准确率82.6%。
- 效果接近显式帧级标注模型,适合解释性研究与弱监督场景。
基于梯度的归因方法可突出神经网络预测中重要的输入区域,但其在音频分类中对时序声音事件检测的有效性尚未系统评估。本文评估积分梯度(IG)在未接受时序监督训练的分类器上,是否能准确识别声音事件的时间边界。我们使用带有真实时间戳的合成多声源音频进行测试。在10类家庭声音数据集上,IG实现平均交并比(IoU)0.39、帧级F1分数0.52、点位游戏准确率82.6%。对比基线:弱监督帧级卷积网络(FW-WS)达到0.42 IoU、0.55 F1、97.3% PG;强监督版本(FW-SS)达0.45 IoU、0.58 F1、97.9% PG。结果表明,后处理的IG能捕捉有意义的声音事件时序活动模式,定位性能接近显式生成帧级预测的模型。所有方法显著优于随机和能量基基线。
原文摘要 · Abstract (English)
Gradient-based attribution methods can highlight input regions important for neural network predictions, but their effectiveness for temporal sound event detection in audio classification has not been systematically evaluated. This paper assesses whether integrated gradients (IG) can temporally detect sound events when applied to a classifier trained without temporal supervision. We use synthetic polyphonic audio with ground truth timestamps to measure alignment between IG attributions and event boundaries. On a 10-class domestic sound dataset, IG achieves mean Intersection over Union (IoU) of 0.39, frame-level F1 of 0.52, and Pointing Game accuracy of 82.6\%. For comparison, a framewise CNN trained with weak supervision (FW-WS, clip-level training labels) achieves 0.42 IoU, 0.55 F1, and 97.3\% PG, while a strongly supervised variant (FW-SS, frame-level training labels) reaches 0.45 IoU, 0.58 F1, and 97.9\% PG. Overall, these results suggest that post-hoc IG captures meaningful temporal activity patterns of sound events, with localization performance approaching models that explicitly produce frame-level predictions. All methods substantially outperform random and energy-based baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。