arXiv:2411.19806cs.SDcs.AI2024-11中稿 · the IEEE Internati…被引 1

用联合嵌入架构实现零样本音乐分轨检索,无需训练即可匹配新乐器。

Zero-shot Musical Stem Retrieval with Joint-Embedding Predictive Architectures

  • 通过联合训练编码器与预测器,生成上下文与目标的潜在表示。
  • 在MUSDB18和MoisesDB上显著优于基线,支持未见乐器的零样本检索。
  • 预训练编码器用对比学习,提升性能;嵌入保留时间结构,适合下游任务。

本文研究音乐混合音中分轨的零样本检索问题:给定一段混音,目标是找出能自然融合的分轨。提出基于联合嵌入预测架构的新方法,编码器与预测器联合训练,生成上下文与目标的潜在表示。特别地,预测器设计为可接受任意乐器作为条件,使模型具备零样本检索能力。实验表明,使用对比学习预训练编码器可显著提升性能。在MUSDB18和MoisesDB数据集上的结果表明,该模型优于现有基线,能支持更精确或未见的条件输入。此外,在节拍跟踪任务上评估了学习到的嵌入,验证其保留了时间结构与局部信息。

原文摘要 · Abstract (English)

In this paper, we tackle the task of musical stem retrieval. Given a musical mix, it consists in retrieving a stem that would fit with it, i.e., that would sound pleasant if played together. To do so, we introduce a new method based on Joint-Embedding Predictive Architectures, where an encoder and a predictor are jointly trained to produce latent representations of a context and predict latent representations of a target. In particular, we design our predictor to be conditioned on arbitrary instruments, enabling our model to perform zero-shot stem retrieval. In addition, we discover that pretraining the encoder using contrastive learning drastically improves the model's performance. We validate the retrieval performances of our model using the MUSDB18 and MoisesDB datasets. We show that it significantly outperforms previous baselines on both datasets, showcasing its ability to support more or less precise (and possibly unseen) conditioning. We also evaluate the learned embeddings on a beat tracking task, demonstrating that they retain temporal structure and local information.

音乐生成零样本嵌入模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。