对比6种预训练音频模型在音乐推荐中的表现差异。
Comparative Analysis of Pretrained Audio Representations in Music Recommender Systems
- 用6个预训练模型提取音频特征,适配三种推荐算法
- 不同模型在推荐任务中表现差异显著,非通用有效
- 为提升音乐推荐系统提供可复用的音频表征研究基础
近年来,音乐信息检索(MIR)提出了多种在大规模音乐数据上预训练的模型。迁移学习已证明预训练后端模型在众多下游任务(如自动标注和流派分类)中的有效性。然而,现有MIR研究普遍未探讨这些模型在音乐推荐系统(MRS)中的效率。同时,推荐系统领域更倾向使用传统端到端神经网络学习方法。本研究填补这一空白,评估六种预训练后端模型(MusicFM、Music2Vec、MERT、EncodecMAE、Jukebox、MusiCNN)在MRS中的适用性。采用三种推荐模型:K近邻(KNN)、浅层神经网络和BERT4Rec进行性能评估。结果表明,预训练音频表示在传统MIR任务与MRS之间的表现存在显著差异,说明后端模型捕捉的音乐信息价值随任务而异。本研究为未来探索预训练音频表示以提升音乐推荐系统奠定了基础。
原文摘要 · Abstract (English)
Over the years, Music Information Retrieval (MIR) has proposed various models pretrained on large amounts of music data. Transfer learning showcases the proven effectiveness of pretrained backend models with a broad spectrum of downstream tasks, including auto-tagging and genre classification. However, MIR papers generally do not explore the efficiency of pretrained models for Music Recommender Systems (MRS). In addition, the Recommender Systems community tends to favour traditional end-to-end neural network learning over these models. Our research addresses this gap and evaluates the applicability of six pretrained backend models (MusicFM, Music2Vec, MERT, EncodecMAE, Jukebox, and MusiCNN) in the context of MRS. We assess their performance using three recommendation models: K-nearest neighbours (KNN), shallow neural network, and BERT4Rec. Our findings suggest that pretrained audio representations exhibit significant performance variability between traditional MIR tasks and MRS, indicating that valuable aspects of musical information captured by backend models may differ depending on the task. This study establishes a foundation for further exploration of pretrained audio representations to enhance music recommendation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。