用9种预训练音频模型提升音乐推荐,验证其在冷热启动场景下的有效性。
Adopting State-of-the-Art Pretrained Audio Representations for Music Recommender Systems

- 选用9种预训练音频模型,结合5种推荐方法评估性能。
- 预训练模型在冷启动推荐中表现显著优于传统方法。
- 揭示音频特征对推荐任务的适配性差异,适合音乐推荐研究者。
近年来,音乐信息检索(MIR)领域发布了多种在大量音乐数据上预训练的模型。迁移学习已证明预训练后端模型在自动标记、流派分类等下游任务中的有效性。然而,现有MIR研究普遍未探索预训练模型在音乐推荐系统(MRS)中的效率。同时,推荐系统领域更倾向于采用传统的端到端神经网络训练。本研究填补这一空白,评估了九种预训练后端模型(MusicFM、Music2Vec、MERT、EncodecMAE、Jukebox、MusiCNN、MULE、MuQ和MuQ-MuLan)在MRS中的表现。我们采用五种推荐方法:K近邻(KNN)、浅层神经网络、对比多模态投影、混合模型以及BERT4Rec,分别在热启动和冷启动场景下进行测试。结果表明,预训练音频表示在传统MIR任务与音乐推荐任务之间存在显著性能差异,说明后端模型所捕捉的音乐信息在不同任务中具有不同价值。该研究为利用预训练音频表示改进音乐推荐系统奠定了基础。
原文摘要 · Abstract (English)
Over the years, Music Information Retrieval (MIR) research community has released various models pretrained on large amounts of music data. Transfer learning showcases the proven effectiveness of pretrained backend models for a broad spectrum of downstream tasks, including auto-tagging and genre classification. However, MIR papers generally do not explore the efficiency of pretrained models for Music Recommender Systems (MRS). In addition, the Recommender Systems community tends to favour traditional end-to-end neural network training. Our research addresses this gap and evaluates the performance of nine pretrained backend models (MusicFM, Music2Vec, MERT, EncodecMAE, Jukebox, MusiCNN, MULE, MuQ and MuQ-MuLan) in the context of MRS. We assess them using five recommendation approaches: K-Nearest Neighbours (KNN), Shallow Neural Network, Contrastive Multi-Modal projection, a Hybrid model, and BERT4Rec both for the hot and cold-start scenarios. Our findings suggest that pretrained audio representations exhibit significant performance disparity between traditional MIR tasks and both hot and cold music recommendations, indicating that valuable aspects of musical information captured by backend models may differ depending on the task. This study establishes a foundation for further exploration of pretrained audio representations to enhance music recommendation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。