arXiv:2409.06196cs.SDcs.LG2024-09

提出双分支架构,提升异构声音事件检测性能

MTDA-HSED: Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event Detection

  • 设计互助适配器与双分支融合模块,应对多场景和多粒度挑战
  • 在DESED和MAESTRO Real数据集上,mpAUC指标提升5%
  • 适合音频理解、声学场景分析领域的研究者参考

声音事件检测(SED)在理解与感知声学场景中起关键作用。现有方法虽表现优异,但在异构数据集上学习复杂场景特征的能力不足。本文提出一种新型双分支架构——用于异构声音事件检测的互助调优与双分支聚合(MTDA-HSED)。该架构采用互助音频适配器(M3A),嵌入BEATs模块作为适配器,在多场景数据集上微调以提升性能;同时引入双分支中层融合(DBMF)模块,连接BEATs与CNN分支,实现深度信息融合。实验表明,所提方法在DESED和MAESTRO Real数据集上,相较于基线模型,mpAUC指标提升5%。代码已开源。

原文摘要 · Abstract (English)

Sound Event Detection (SED) plays a vital role in comprehending and perceiving acoustic scenes. Previous methods have demonstrated impressive capabilities. However, they are deficient in learning features of complex scenes from heterogeneous dataset. In this paper, we introduce a novel dual-branch architecture named Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event Detection (MTDA-HSED). The MTDA-HSED architecture employs the Mutual-Assistance Audio Adapter (M3A) to effectively tackle the multi-scenario problem and uses the Dual-Branch Mid-Fusion (DBMF) module to tackle the multi-granularity problem. Specifically, M3A is integrated into the BEATs block as an adapter to improve the BEATs' performance by fine-tuning it on the multi-scenario dataset. The DBMF module connects BEATs and CNN branches, which facilitates the deep fusion of information from the BEATs and the CNN branches. Experimental results show that the proposed methods exceed the baseline of mpAUC by \textbf{$5\%$} on the DESED and MAESTRO Real datasets. Code is available at https://github.com/Visitor-W/MTDA.

声音检测双分支音频处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。