用合成数据预训练声音事件定位检测模型,提升低资源场景下的性能
PSELDNets: Pre-trained Neural Networks on a Large-scale Synthetic Dataset for Sound Event Localization and Detection
- 基于170类声音的1167小时合成音频预训练模型
- 在真实数据集上超越现有最先进方法,尤其在少样本时表现优异
- 适配器比特技术让模型可用单声道音频高效迁移,适合实际部署
声音事件定位与检测(SELD)近年来通过学习方法取得显著进展。这类系统通常在特定数据集上从零开始训练,具备一定泛化能力。近期大规模数据预训练的声学事件分类(SEC)模型取得突破,引发一个关键问题:能否将此类进展扩展至构建SELD基础模型?本文提出在大规模合成数据集上预训练的SELD网络(PSELDNets)。该合成数据集通过卷积声源与模拟空间混响响应(SRIRs)生成,包含1,167小时音频片段,涵盖170种声音类别。在多种SELD场景中应用时,针对低资源情况,引入一种数据高效的微调方法AdapterBit。PSELDNets在基于TAU-SRIR DB收集的混响响应的合成测试集上表现良好,并验证了其在三个公开数据集及自建真实录音上的可迁移性。结果表明,所有公开数据集上均超越当前最优系统。传统SELD依赖充足多通道音频,而引入AdapterBit后,仅需少量或多通道甚至单声道音频即可实现高效适应,优于传统微调方式。
原文摘要 · Abstract (English)
Sound event localization and detection (SELD) has seen substantial advancements through learning-based methods. These systems, typically trained from scratch on specific datasets, have shown considerable generalization capabilities. Recently, deep neural networks trained on large-scale datasets have achieved remarkable success in the sound event classification (SEC) field, prompting an open question of whether these advances can be extended to the development of SELD foundation models. In this paper, leveraging the power of pre-trained SEC models, we propose pre-trained SELD networks (PSELDNets) on a large-scale synthetic dataset. The synthetic dataset, generated by convolving sound events with simulated spatial room impulse responses (SRIRs), contains 1,167 hours of audio clips with an ontology of 170 sound classes. These PSELDNets are applied to various SELD scenarios. When we adapt PSELDNets to specific scenarios, particularly in cases of low-resource data, we introduce a data-efficient fine-tuning method, AdapterBit. PSELDNets are evaluated on synthetic-test-set using collected SRIRs from the TAU Spatial Room Impulse Response Database (TAU-SRIR DB) and achieve satisfactory performance. We also carried out experiments to validate the transferability of PSELDNets to three publicly available datasets and our own real-world recordings. The results demonstrate that PSELDNets surpass state-of-the-art systems across all publicly available datasets. Given the need for direction-of-arrival estimation, SELD generally relies on sufficient multi-channel audio clips. However, incorporating the AdapterBit, PSELDNets show more efficient adaptability to various scenarios using minimal multi-channel or even just monophonic audio clips, outperforming traditional fine-tuning approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。