让自监督音频模型更好处理重叠声音,提升真实场景下的表现
SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes
- 用混合音频训练自监督模型,增强对复杂声音的感知能力
- 在多音源数据集上性能提升最高达9.1%(mAP),同时保持单音源优势
- 适合需要高鲁棒性音频理解的应用,如智能音箱、环境监控
自监督预训练音频模型广泛应用于实际系统,尤其在多模态大语言模型中。这些模型通常以冻结状态使用,假设预训练已使其具备处理真实音频的能力。然而,现实音频常为多音源重叠的复杂声景,而现有评测多基于单音源数据集(如环境音、语音),导致模型在真实场景下的泛化能力未被充分检验。为此,本文提出自监督音频混合学习(SSLAM),旨在提升模型从多音源数据中学习的能力,同时维持对单音源任务的优异表现。我们在主流单音源基准上评估SSLAM,并与多种先进方法在多个高质量公开多音源数据集上进行对比。结果表明,SSLAM不仅在多音源任务上实现显著提升,最大达到9.1%(mAP);在标准音频SSL基准上也保持或超越现有水平,其中AudioSet-2M(AS-2M)上达到50.2%的mAP,提升3.9%。
原文摘要 · Abstract (English)
Self-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed in a frozen state, under the assumption that the SSL pre-training has sufficiently equipped them to handle real-world audio. However, a critical question remains: how well do these models actually perform in real-world conditions, where audio is typically polyphonic and complex, involving multiple overlapping sound sources? Current audio SSL methods are often benchmarked on datasets predominantly featuring monophonic audio, such as environmental sounds, and speech. As a result, the ability of SSL models to generalize to polyphonic audio, a common characteristic in natural scenarios, remains underexplored. This limitation raises concerns about the practical robustness of SSL models in more realistic audio settings. To address this gap, we introduce Self-Supervised Learning from Audio Mixtures (SSLAM), a novel direction in audio SSL research, designed to improve, designed to improve the model's ability to learn from polyphonic data while maintaining strong performance on monophonic data. We thoroughly evaluate SSLAM on standard audio SSL benchmark datasets which are predominantly monophonic and conduct a comprehensive comparative analysis against SOTA methods using a range of high-quality, publicly available polyphonic datasets. SSLAM not only improves model performance on polyphonic audio, but also maintains or exceeds performance on standard audio SSL benchmarks. Notably, it achieves up to a 3.9\% improvement on the AudioSet-2M (AS-2M), reaching a mean average precision (mAP) of 50.2. For polyphonic datasets, SSLAM sets new SOTA in both linear evaluation and fine-tuning regimes with performance improvements of up to 9.1\% (mAP).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。