对比多种自监督模型,提升声音事件检测的融合与后处理效果
Revisiting SSL for sound event detection: complementary fusion and adaptive post-processing
- 融合不同自监督模型特征,采用双模态策略增强性能
- 引入归一化边界框后处理,使独立模型性能提升4%
- 实验证明多模型互补性,指导特定任务的融合设计
自监督学习(SSL)模型为声音事件检测(SED)提供了强大表征,但其协同潜力尚未充分挖掘。本研究系统评估了当前最先进的SSL模型,以指导SED中最优模型的选择与集成。提出一个框架,通过三种融合策略结合异构SSL表征(如BEATs、HuBERT、WavLM):单一嵌入融合、双模态融合和完全聚合。在DCASE 2023 Task 4挑战数据集上的实验表明,双模态融合(如CRNN+BEATs+WavLM)带来互补性能提升,而仅使用CRNN+BEATs在单个SSL模型中表现最佳。此外,提出归一化声音事件边界框(nSEBBs),一种自适应后处理方法,可动态调整事件边界预测,使独立SSL模型的PSDS1最高提升4%。结果表明SSL架构具有兼容性与互补性,为任务特定融合与鲁棒SED系统设计提供依据。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) models offer powerful representations for sound event detection (SED), yet their synergistic potential remains underexplored. This study systematically evaluates state-of-the-art SSL models to guide optimal model selection and integration for SED. We propose a framework that combines heterogeneous SSL representations (e.g., BEATs, HuBERT, WavLM) through three fusion strategies: individual SSL embedding integration, dual-modal fusion, and full aggregation. Experiments on the DCASE 2023 Task 4 Challenge reveal that dual-modal fusion (e.g., CRNN+BEATs+WavLM) achieves complementary performance gains, while CRNN+BEATs alone delivers the best results among individual SSL models. We further introduce normalized sound event bounding boxes (nSEBBs), an adaptive post-processing method that dynamically adjusts event boundary predictions, improving PSDS1 by up to 4% for standalone SSL models. These findings highlight the compatibility and complementarity of SSL architectures, providing guidance for task-specific fusion and robust SED system design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。