用大模型评估推荐系统中的意外之喜,更准更通用。
A Universal Framework for Offline Serendipity Evaluation in Recommender Systems via Large Language Models
- 用大模型做评价器,通过思维链提示提升判断准确率
- 三组真实数据测试显示,专用推荐系统不总胜出
- 适合想客观评估推荐惊喜度的研究者和工程师
推荐系统中的意外之喜(serendipity)能提升用户满意度,但其效果难以评估,因真实情况通常不可观测。现有离线指标常依赖模糊定义或仅适用于特定数据集与推荐系统,泛化能力差。为此,我们提出一个通用评估框架,利用具备广泛知识与推理能力的大语言模型(LLMs)作为评价者。首先,在包含用户标注意外之喜真实标签的数据集上,我们测试了四种不同提示策略下LLM的预测准确率,发现思维链提示(chain-of-thought prompt)表现最佳。随后,在三个常用真实世界数据集上,无真实标签条件下,我们使用该框架重新评估了专门设计的意外之喜推荐系统与通用推荐系统的性能。结果表明,没有一个专门的意外之喜推荐系统在所有数据集上都持续领先,甚至通用推荐系统有时表现更好。
原文摘要 · Abstract (English)
Serendipity in recommender systems (RSs) has attracted increasing attention as a concept that enhances user satisfaction by presenting unexpected and useful items. However, evaluating serendipitous performance remains challenging because its ground truth is generally unobservable. The existing offline metrics often depend on ambiguous definitions or are tailored to specific datasets and RSs, thereby limiting their generalizability. To address this issue, we propose a universally applicable evaluation framework that leverages large language models (LLMs) known for their extensive knowledge and reasoning capabilities, as evaluators. First, to improve the evaluation performance of the proposed framework, we assessed the serendipity prediction accuracy of LLMs using four different prompt strategies on a dataset containing user-annotated serendipitous ground truth and found that the chain-of-thought prompt achieved the highest accuracy. Next, we re-evaluated the serendipitous performance of both serendipity-oriented and general RSs using the proposed framework on three commonly used real-world datasets, without the ground truth. The results indicated that there was no serendipity-oriented RS that consistently outperformed across all datasets, and even a general RS sometimes achieved higher performance than the serendipity-oriented RS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。