提出LOEV框架,让音乐表示模型保留关键信息,提升任务适应性。
Leave-One-EquiVariant: Alleviating invariance-related information loss in contrastive music representations
- 通过选择性保留特定数据增强信息,避免过度不变性导致的信息丢失。
- 在多种音乐检索任务中表现更优,且不降低通用表示质量。
- 适合需要关注音乐属性变化的下游任务,如风格识别、节奏分析。
对比学习在自监督音乐表征学习中表现优异,尤其在音乐信息检索(MIR)任务中。然而,依赖增强链生成对比视图及由此产生的不变性,会带来不同下游任务对某些音乐属性敏感性不足的问题。为此,我们提出Leave One EquiVariant(LOEV)框架,相较以往方法更具灵活性与任务适应性,能选择性保留特定增强相关的信息,使模型维持任务相关的等变性。实验表明,LOEV有效缓解了因学习到的不变性带来的信息损失,在增强相关任务和检索任务中均取得性能提升,同时保持良好的通用表征能力。此外,我们还提出LOEV++,通过自监督方式构建解耦的潜在空间,实现基于增强相关属性的精准检索。
原文摘要 · Abstract (English)
Contrastive learning has proven effective in self-supervised musical representation learning, particularly for Music Information Retrieval (MIR) tasks. However, reliance on augmentation chains for contrastive view generation and the resulting learnt invariances pose challenges when different downstream tasks require sensitivity to certain musical attributes. To address this, we propose the Leave One EquiVariant (LOEV) framework, which introduces a flexible, task-adaptive approach compared to previous work by selectively preserving information about specific augmentations, allowing the model to maintain task-relevant equivariances. We demonstrate that LOEV alleviates information loss related to learned invariances, improving performance on augmentation related tasks and retrieval without sacrificing general representation quality. Furthermore, we introduce a variant of LOEV, LOEV++, which builds a disentangled latent space by design in a self-supervised manner, and enables targeted retrieval based on augmentation related attributes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。