通过联合预测掩码隐变量和无监督分类,提升音频自监督表征性能。
Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning
- 双任务训练:掩码隐变量预测 + 教师-学生分布匹配
- 在OpenMIC等4个数据集上达到最优效果
- 适合音频分类与音乐自动标记任务
近期基于掩码隐变量预测的自监督学习方法已证明能生成强大的输入表征。然而,在训练过程中,学习到的隐空间可进一步转换以提取更适合作下游分类任务的高级信息。为此,我们提出一种新方法:掩码隐变量预测与分类(MATPAC),通过联合求解两个预训练任务进行训练。第一项任务为掩码隐变量预测,确保隐空间中具备鲁棒的输入表征;第二项为无监督分类,利用第一项任务的隐表示来匹配教师模型与学生模型之间的概率分布。我们在多个基准音频分类数据集(如OpenMIC、GTZAN、ESC-50、US8K)上验证了MATPAC方法,并进行了消融实验。结果表明,MATPAC在这些数据集上达到当前最佳自监督学习性能,且在Magna-tag-a-tune数据集上的音乐自动标记任务中超越了同类有监督方法。
原文摘要 · Abstract (English)
Recently, self-supervised learning methods based on masked latent prediction have proven to encode input data into powerful representations. However, during training, the learned latent space can be further transformed to extract higher-level information that could be more suited for downstream classification tasks. Therefore, we propose a new method: MAsked latenT Prediction And Classification (MATPAC), which is trained with two pretext tasks solved jointly. As in previous work, the first pretext task is a masked latent prediction task, ensuring a robust input representation in the latent space. The second one is unsupervised classification, which utilises the latent representations of the first pretext task to match probability distributions between a teacher and a student. We validate the MATPAC method by comparing it to other state-of-the-art proposals and conducting ablations studies. MATPAC reaches state-of-the-art self-supervised learning results on reference audio classification datasets such as OpenMIC, GTZAN, ESC-50 and US8K and outperforms comparable supervised methods results for musical auto-tagging on Magna-tag-a-tune.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。