首个评估大模型智能体推荐系统的基准测试,验证其在个性化推荐中的优势。
AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems
- 构建交互式文本推荐模拟器,支持多种推荐场景。
- 对比10种经典与智能体推荐方法,证明智能体系统更优。
- 提供开源平台和持续排行榜,适合研究者复现与改进。
基于大语言模型(LLMs)的智能体推荐系统代表了个性化推荐领域的一次范式转变,利用LLMs的高级推理与角色扮演能力,实现自主、自适应决策。与传统推荐方法不同,智能体系统可动态获取并解析复杂环境中的用户-物品交互,生成泛化能力强的推荐策略。然而,该领域尚缺乏标准化评估协议。为此,我们提出:(1) 一个包含丰富用户与物品元数据的交互式文本推荐模拟器,涵盖经典、兴趣演化和冷启动三种典型评估场景;(2) 一个统一的模块化框架,用于开发与研究智能体推荐系统;(3) 首个全面基准,对比10种经典与智能体推荐方法。实验结果表明智能体系统具有显著优势,并为关键组件设计提供了可操作指南。该基准环境已通过公开挑战赛验证,持续维护排行榜,支持社区协作与可复现研究。基准数据集已公开:https://huggingface.co/datasets/SGJQovo/AgentRecBench。
原文摘要 · Abstract (English)
The emergence of agentic recommender systems powered by Large Language Models (LLMs) represents a paradigm shift in personalized recommendations, leveraging LLMs' advanced reasoning and role-playing capabilities to enable autonomous, adaptive decision-making. Unlike traditional recommendation approaches, agentic recommender systems can dynamically gather and interpret user-item interactions from complex environments, generating robust recommendation strategies that generalize across diverse scenarios. However, the field currently lacks standardized evaluation protocols to systematically assess these methods. To address this critical gap, we propose: (1) an interactive textual recommendation simulator incorporating rich user and item metadata and three typical evaluation scenarios (classic, evolving-interest, and cold-start recommendation tasks); (2) a unified modular framework for developing and studying agentic recommender systems; and (3) the first comprehensive benchmark comparing 10 classical and agentic recommendation methods. Our findings demonstrate the superiority of agentic systems and establish actionable design guidelines for their core components. The benchmark environment has been rigorously validated through an open challenge and remains publicly available with a continuously maintained leaderboard~\footnote[2]{https://tsinghua-fib-lab.github.io/AgentSocietyChallenge/pages/overview.html}, fostering ongoing community engagement and reproducible research. The benchmark is available at: \hyperlink{https://huggingface.co/datasets/SGJQovo/AgentRecBench}{https://huggingface.co/datasets/SGJQovo/AgentRecBench}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。