arXiv:2601.01997cs.IRcs.AI2026-01

对比ChatGPT-3.5与4在推荐多样性、新颖性及流行度偏见上的表现

Exploring Diversity, Novelty, and Popularity Bias in ChatGPT's Recommendations

  • 用三组数据集评估模型在Top-N和冷启动场景下的推荐能力
  • ChatGPT-4在新颖性和多样性上优于传统推荐系统,冷启动时更胜一筹
  • 适合关注推荐系统长期体验与新用户适配的研究者参考

ChatGPT已成为跨领域通用工具,其在推荐系统中的应用日益受到关注,但现有研究多聚焦于准确率。本研究首次系统评估ChatGPT-3.5与ChatGPT-4在推荐多样性、新颖性及流行度偏见方面的表现。基于三个独立数据集,在Top-N推荐与冷启动场景下进行测试。结果表明,ChatGPT-4在多样性和新颖性方面达到或超越传统推荐模型;尤其在冷启动场景中,不仅准确率更高,且推荐内容更具新颖性。该研究揭示了大模型在推荐系统中的潜力与局限,为超越精度指标的个性化推荐提供了新视角。

原文摘要 · Abstract (English)

ChatGPT has emerged as a versatile tool, demonstrating capabilities across diverse domains. Given these successes, the Recommender Systems (RSs) community has begun investigating its applications within recommendation scenarios primarily focusing on accuracy. While the integration of ChatGPT into RSs has garnered significant attention, a comprehensive analysis of its performance across various dimensions remains largely unexplored. Specifically, the capabilities of providing diverse and novel recommendations or exploring potential biases such as popularity bias have not been thoroughly examined. As the use of these models continues to expand, understanding these aspects is crucial for enhancing user satisfaction and achieving long-term personalization. This study investigates the recommendations provided by ChatGPT-3.5 and ChatGPT-4 by assessing ChatGPT's capabilities in terms of diversity, novelty, and popularity bias. We evaluate these models on three distinct datasets and assess their performance in Top-N recommendation and cold-start scenarios. The findings reveal that ChatGPT-4 matches or surpasses traditional recommenders, demonstrating the ability to balance novelty and diversity in recommendations. Furthermore, in the cold-start scenario, ChatGPT models exhibit superior performance in both accuracy and novelty, suggesting they can be particularly beneficial for new users. This research highlights the strengths and limitations of ChatGPT's recommendations, offering new perspectives on the capacity of these models to provide recommendations beyond accuracy-focused metrics.

推荐系统大模型冷启动多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。