arXiv:2409.09026cs.SDcs.AI2024-09中稿 · the 2nd Music Reco…被引 2

用对比预训练音频嵌入提升音乐推荐效果

Towards Leveraging Contrastively Pretrained Neural Audio Embeddings for Recommender Tasks

  • 用CLAP等对比预训练模型提取音乐深层特征
  • 在图框架下显著改善冷启动场景推荐性能
  • 适合做音乐推荐系统且关注内容特征的研究者

音乐推荐系统常使用基于网络的模型捕捉曲目、艺术家与用户间的关系。尽管这些关系能提供有价值的预测信息,但新曲目或新艺术家因初始信息不足而面临冷启动问题。为解决此问题,可直接从音乐中提取内容特征以增强基于协同过滤的方法。以往方法依赖手工设计的音频特征,本文探索使用对比预训练的神经音频嵌入模型,其能提供更丰富、更细致的音乐表示。实验表明,尤其是使用对比语言-音频预训练(CLAP)模型生成的嵌入,在图结构框架中显著提升了音乐推荐任务的表现。

原文摘要 · Abstract (English)

Music recommender systems frequently utilize network-based models to capture relationships between music pieces, artists, and users. Although these relationships provide valuable insights for predictions, new music pieces or artists often face the cold-start problem due to insufficient initial information. To address this, one can extract content-based information directly from the music to enhance collaborative-filtering-based methods. While previous approaches have relied on hand-crafted audio features for this purpose, we explore the use of contrastively pretrained neural audio embedding models, which offer a richer and more nuanced representation of music. Our experiments demonstrate that neural embeddings, particularly those generated with the Contrastive Language-Audio Pretraining (CLAP) model, present a promising approach to enhancing music recommendation tasks within graph-based frameworks.

音乐推荐音频嵌入图神经网络冷启动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。