用原型对比学习提升少样本音频分类效果
Prototypical Contrastive Learning For Improved Few-Shot Audio Classification
- 结合监督对比损失与原型学习,优化音频特征表示
- 在5类5样本设置下达到当前最好性能
- 适合研究少样本音频识别的学者参考
少样本学习已成为在标注数据有限情况下训练模型的强大范式,尤其适用于大规模标注不现实的场景。尽管图像领域已有大量研究,但音频分类中的少样本学习仍相对薄弱。本文探讨将监督对比损失融入原型少样本训练对音频分类的影响。具体而言,我们证明角距离损失相比标准对比损失可进一步提升性能。方法采用SpecAugment配合自注意力机制,将多种增强输入版本的信息融合为统一嵌入表示。在MetaAudio基准上进行评估,该基准包含五个数据集,具有预定义划分、标准化预处理及丰富的少样本模型对比基线。所提方法在5类5样本设置下实现当前最优性能。
原文摘要 · Abstract (English)
Few-shot learning has emerged as a powerful paradigm for training models with limited labeled data, addressing challenges in scenarios where large-scale annotation is impractical. While extensive research has been conducted in the image domain, few-shot learning in audio classification remains relatively underexplored. In this work, we investigate the effect of integrating supervised contrastive loss into prototypical few shot training for audio classification. In detail, we demonstrate that angular loss further improves the performance compared to the standard contrastive loss. Our method leverages SpecAugment followed by a self-attention mechanism to encapsulate diverse information of augmented input versions into one unified embedding. We evaluate our approach on MetaAudio, a benchmark including five datasets with predefined splits, standardized preprocessing, and a comprehensive set of few-shot learning models for comparison. The proposed approach achieves state-of-the-art performance in a 5-way, 5-shot setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。