LLM让音乐推荐从打分转向对话,但评估标准需重想。
Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation
- 用自然语言交互重构音乐推荐逻辑,突破传统打分模式。
- 提出评估新维度:效果、风险与可解释性三者并重。
- 适合关注AI评价体系变革的推荐系统研究者。
音乐推荐系统长期依赖信息检索范式,以召回任务准确率衡量进展。然而这种简化方法难以回答何为优质推荐的本质问题。尽管已有用户研究或公平性分析尝试拓宽评估,但影响有限。大语言模型(LLMs)的出现打破了这一框架:其生成式特性使传统准确率指标失效,同时引入幻觉、知识截止、非确定性及训练数据不透明等挑战,使常规训练/测试流程难以解读。另一方面,LLMs也带来新机遇,支持自然语言交互,甚至使模型能充当评估者。本文主张,向LLM驱动的音乐推荐系统转型,需重新思考评估范式。我们首先回顾LLMs在用户建模、物品建模及基于自然语言的推荐中的影响;接着分析NLP领域的评估方法,提炼对音乐推荐适用的实践与开放问题;最后综合观点,聚焦于提示工程在音乐推荐中的应用,提出一套结构化的成功与风险维度。目标是为音乐推荐社区提供一个更新的、教学性的、跨学科的评估视角。
原文摘要 · Abstract (English)
Music Recommender Systems (MRSs) have long relied on an information retrieval framing, where progress is measured mainly through accuracy on retrieval-oriented subtasks. While effective, this reductionist paradigm struggles to address the deeper question of what makes a good recommendation. Attempts to broaden evaluation, through user studies or fairness analyses, have had limited impact. The emergence of Large Language Models (LLMs) disrupts this framework: LLMs are generative rather than ranking-based, making standard accuracy metrics questionable. They also introduce challenges such as hallucinations, knowledge cutoffs, non-determinism, and opaque training data, rendering traditional train/test protocols difficult to interpret. At the same time, LLMs create new opportunities, enabling natural language (NL) interaction and even allowing models to act as evaluators. This work argues that the shift toward LLM-driven MRSs requires rethinking evaluation. We first review how LLMs have impacted user modeling, item modeling, and NL-based recommendation in music. We then analyze evaluation practices from NLP, highlighting methodologies and open challenges relevant to MRSs. Finally, we synthesize insights, focusing on how LLM prompting applies to MRSs, to outline a structured set of success and risk dimensions. Our goal is to provide the MRSs community with an updated, pedagogical, and cross-disciplinary perspective on evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。