提升音频文本模型在复杂背景下的声音分类准确率
Domain Adaptation Method and Modality Gap Impact in Audio-Text Models for Prototypical Sound Classification
- 引入背景声源贡献度量化方法,无需重训练即可增强鲁棒性
- 在低信噪比条件下性能显著提升,跨背景泛化能力更强
- 适合需要零样本声音分类的智能设备与环境监测场景
音频-文本模型广泛应用于零样本环境声音分类,可减少对标注数据的依赖。然而,我们发现当存在背景音源时,其性能严重下降。分析表明,这种退化主要由背景音景的信噪比(SNR)决定,与背景类型无关。为此,我们提出一种新方法,量化并整合背景声源对分类的贡献,提升性能且无需模型重训练。该领域自适应技术在多种背景和信噪比条件下均有效提升准确率。此外,我们分析了音频与文本嵌入间的模态差距,发现缩小此差距可改善分类表现。该方法在多种前沿原型学习框架中均具良好泛化性,展现出强可扩展性与鲁棒性。
原文摘要 · Abstract (English)
Audio-text models are widely used in zero-shot environmental sound classification as they alleviate the need for annotated data. However, we show that their performance severely drops in the presence of background sound sources. Our analysis reveals that this degradation is primarily driven by SNR levels of background soundscapes, and independent of background type. To address this, we propose a novel method that quantifies and integrates the contribution of background sources into the classification process, improving performance without requiring model retraining. Our domain adaptation technique enhances accuracy across various backgrounds and SNR conditions. Moreover, we analyze the modality gap between audio and text embeddings, showing that narrowing this gap improves classification performance. The method generalizes effectively across state-of-the-art prototypical approaches, showcasing its scalability and robustness for diverse environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。