arXiv:2507.17290cs.IR2025-07

用大模型模拟用户评估推荐系统的意外之喜,效果比传统方法更好。

Exploring the Potential of LLMs for Serendipity Evaluation in Recommender Systems

  • 让大模型直接模拟人类用户判断推荐结果是否意外惊喜
  • 零样本大模型性能已达或超越传统指标,多模型协作更优
  • 适合做推荐系统评测的学者和工程师快速验证新算法

意外之喜在推荐系统中对提升用户满意度至关重要,但其评价因主观性强、概念模糊而困难。现有算法多依赖代理指标间接评估,常与真实用户感知脱节。随着大语言模型(LLMs)在各类人工标注任务中革新评价方法,我们探索核心问题:大模型能否有效模拟人类用户进行意外之喜评估?我们在电商与电影两个真实用户研究数据集上开展元评估,重点分析三方面:大模型相比传统代理指标的准确性、辅助数据对大模型理解的影响、近期流行的多大模型技术的有效性。结果显示,即使最简单的零样本大模型也达到或超过传统指标表现;多大模型协同与引入辅助数据可进一步提升与人类观点的一致性。最优大模型评估方案与用户研究结果的皮尔逊相关系数达21.5%。研究表明,大模型可能成为准确且低成本的评估工具,为推荐系统中的意外之喜评价带来新范式。

原文摘要 · Abstract (English)

Serendipity plays a pivotal role in enhancing user satisfaction within recommender systems, yet its evaluation poses significant challenges due to its inherently subjective nature and conceptual ambiguity. Current algorithmic approaches predominantly rely on proxy metrics for indirect assessment, often failing to align with real user perceptions, thus creating a gap. With large language models (LLMs) increasingly revolutionizing evaluation methodologies across various human annotation tasks, we are inspired to explore a core research proposition: Can LLMs effectively simulate human users for serendipity evaluation? To address this question, we conduct a meta-evaluation on two datasets derived from real user studies in the e-commerce and movie domains, focusing on three key aspects: the accuracy of LLMs compared to conventional proxy metrics, the influence of auxiliary data on LLM comprehension, and the efficacy of recently popular multi-LLM techniques. Our findings indicate that even the simplest zero-shot LLMs achieve parity with, or surpass, the performance of conventional metrics. Furthermore, multi-LLM techniques and the incorporation of auxiliary data further enhance alignment with human perspectives. Based on our findings, the optimal evaluation by LLMs yields a Pearson correlation coefficient of 21.5\% when compared to the results of the user study. This research implies that LLMs may serve as potentially accurate and cost-effective evaluators, introducing a new paradigm for serendipity evaluation in recommender systems.

推荐系统大模型评估方法意外之喜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。