arXiv:2512.22156cs.SD2025-12被引 5

提出鲁棒框架提升真实场景声音事件定位与检测效果

A Robust framework for sound event localization and detection on real recordings

  • 基于ResNet构建框架,融合真实与模拟数据混合训练
  • 在真实录音上达到领先性能,显著优于基线方法
  • 适合关注真实环境声音分析的研究者与开发者

本技术报告介绍参与DCASE2022挑战赛任务3——声音事件定位与检测(SELD)的系统。该任务旨在识别声音事件的发生及其类别,并估计其空间位置。所提系统采用基于ResNet的模型,在一个新设计的鲁棒框架下实现SELD。为确保在真实声学场景中的泛化能力,框架整合了数据增强技术、真实场景与仿真数据的混合训练流程,以及测试时增强策略。通过引入外部音源和增强技术,模型得以训练出多样化的样本,同时在每个批次中保持真实录音样本的数量,以充分保留真实上下文信息。此外,设计了测试时增强和基于聚类的模型集成方法,用于聚合高置信度预测。实验结果表明,该框架下的模型在真实录音上表现优异,超越基线方法,展现出竞争力。

原文摘要 · Abstract (English)

This technical report describes the systems submitted to the DCASE2022 challenge task 3: sound event localization and detection (SELD). The task aims to detect occurrences of sound events and specify their class, furthermore estimate their position. Our system utilizes a ResNet-based model under a proposed robust framework for SELD. To guarantee the generalized performance on the real-world sound scenes, we design the total framework with augmentation techniques, a pipeline of mixing datasets from real-world sound scenes and emulations, and test time augmentation. Augmentation techniques and exploitation of external sound sources enable training diverse samples and keeping the opportunity to train the real-world context enough by maintaining the number of the real recording samples in the batch. In addition, we design a test time augmentation and a clustering-based model ensemble method to aggregate confident predictions. Experimental results show that the model under a proposed framework outperforms the baseline methods and achieves competitive performance in real-world sound recordings.

声音定位事件检测真实场景深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。