用对比学习稳定训练,结合局部预测提升音乐表征质量。
CoJEPA: Combining Contrastive Learning and JEPA for Global-Local Music Representations
- 共享主干网络同时优化对比和局部预测目标。
- 在音高和和声理解任务上显著优于单一方法。
- 无需额外参数,适合音乐信息检索研究者使用。
联合嵌入预测架构(JEPA)通过潜在空间的自监督预测学习丰富表征,但通常依赖带EMA的教师-学生结构,易产生无信息表征。对比学习训练稳定且生成强全局表征,但在局部任务上受限于其全局目标。本文提出CoJEPA:一个共享主干网络,同时在掩码序列片段上采用JEPA目标,在类别标记上采用对比目标。对比梯度提供训练稳定性,无需EMA教师;而JEPA通过局部预测增强序列标记,这是对比学习无法实现的。关键在于不增加额外参数,仅通过训练信号设计提升表征能力。CoJEPA在全局与局部音乐信息检索任务中均超越或匹配单一方法,尤其在音高与和声理解上优势明显,且无需任务特定结构调整。结果表明,具有互补归纳偏置的目标组合可替代模型规模扩展,鼓励未来关注更智能的训练目标而非更大模型。
原文摘要 · Abstract (English)
Joint-Embedding Predictive Architecture (JEPA) has shown strong performance in learning rich representations through self-supervised prediction in latent space. However, it typically relies on teacher--student architecture with an EMA to stabilise training, and can tend to yield uninformative representations. Contrastive learning is stable to train and produces strong global representations, but remains limited on local tasks by the global nature of its objective. In this work, we combine both into CoJEPA: a single shared backbone jointly trained with a JEPA objective on masked sequence tokens and a contrastive objective on the class token. The contrastive gradient provides stability, removing the need for an EMA teacher entirely, while JEPA enriches the sequence tokens via local predictions that contrastive learning alone cannot provide. Crucially, no extra parameters are added to the backbone: the same model is guided towards richer representations purely through the design of its training signal. CoJEPA takes the best of both worlds, outperforming or matching both individual methods across global and local MIR tasks, with a particularly strong advantage on tonal and harmonic understanding, and without any task-specific architectural changes. CoJEPA shows that combining objectives with complementary inductive biases can substitute for scale, encouraging future work to invest in smarter training objectives over ever-larger models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。