arXiv:2608.06928cs.IR2026-08

对比不同音频模型在音乐推荐中的表现,发现特定预训练模型更适配生成式推荐系统。

From Classification to Recommendation: Empirical Analysis of Audio Embedding Models Application for Content-Based Music Recommendation

论文配图:From Classification to Recommendation: Empirical Analysis of Audio Embedding Models Application for Content-Based Music Recommendation
图 1 · 摘自论文原文
  • 测试六种音频编码器在三类推荐系统中的表现
  • 直接使用预训练嵌入几何结构效果更优,序列训练削弱模型差异
  • 增加语义ID容量未必提升性能,可能引发不稳定

大规模预训练音频表示模型在音频分类与理解任务中表现优异。然而,多数模型优化目标如掩码预测、对比学习或音文对齐,并不必然生成适合推荐系统的表示空间。音乐推荐需捕捉由主观偏好和用户行为塑造的物品关系。尽管音频嵌入已用于传统推荐系统,其在新兴生成式推荐系统中的有效性仍不明确。为此,我们系统评估了六种代表性音频编码器在三类音乐推荐系统(基于内容、序列型、语义ID驱动的生成式)中的表现。进一步研究残差量化设计(包括码本宽度、量化深度及保留的语义ID前缀)对推荐相关性信息保留的影响。在两个音乐推荐数据集上的实验表明:音文对齐与音乐领域预训练表示在直接使用嵌入几何结构时通常更有效;而交互式序列训练显著缩小了编码器间的性能差距。此外,增加语义ID容量并未持续提升生成式推荐系统性能,反而可能导致严重不稳定。这些发现为现代音乐推荐系统选择音频编码器及设计音频衍生语义ID提供了实用指导。

原文摘要 · Abstract (English)

Pretrained audio representation models learned from large-scale corpora have achieved strong performance in audio classification and understanding. However, most existing models are optimized for objectives such as masked prediction, contrastive learning, or audio-text alignment, which do not necessarily produce representation spaces well-suited to recommender systems. Unlike classification, music recommender systems must capture item relationships shaped by subjective and behavior-dependent listener preferences. Although pretrained audio embeddings have been explored in conventional recommender systems, their effectiveness in the rapidly emerging paradigm of generative recommender systems remains underexplored. To address this gap, we systematically evaluate six representative audio encoders across three types of music recommender systems: content-based, sequential, and Semantic-ID-based generative recommender systems. We further investigate how residual-quantization design, including codebook width, quantization depth, and retained Semantic-ID prefixes, affects the preservation of recommendation-relevant information. Experiments on two music recommendation datasets show that audio-text-aligned and music-domain representations are generally more effective when pretrained embedding geometry is used directly, whereas interaction-based sequential training substantially reduces performance differences among encoders. We also find that increasing Semantic-ID capacity does not consistently improve generative recommender systems and may introduce substantial instability. These findings provide practical guidance for selecting audio encoders and designing audio-derived Semantic IDs for modern music recommender systems.

音乐推荐音频嵌入生成式推荐预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。