用大模型代理提升网页平台用户建模与推荐效果
AgentSociety Challenge: Designing LLM Agents for User Modeling and Recommendation on Web Platforms
- 构建跨平台数据+交互模拟器,训练大模型代理理解用户行为
- 开发阶段提升21.9%(用户建模)、20.3%(推荐),最终阶段分别提升9.1%和15.9%
- 开源评测环境,适合研究智能代理与个性化推荐的学者和开发者
AgentSociety Challenge 是首个聚焦大型语言模型(LLM)代理在网页平台用户行为建模与推荐系统中的潜力的竞赛。比赛包含用户建模与推荐两个赛道,参赛者需使用来自 Yelp、Amazon、Goodreads 的联合数据集及交互式环境模拟器,开发创新的 LLM 代理。全球共 295 支队伍参与,累计提交超过 1,400 份作品,历时 37 天。在开发阶段,两个赛道分别实现 21.9% 和 20.3% 的性能提升;决赛阶段分别提升 9.1% 和 15.9%,成果显著。本文详述挑战设计,分析参赛结果,总结最优代理方案。为支持后续研究,我们已开源基准环境:https://tsinghua-fib-lab.github.io/AgentSocietyChallenge。
原文摘要 · Abstract (English)
The AgentSociety Challenge is the first competition in the Web Conference that aims to explore the potential of Large Language Model (LLM) agents in modeling user behavior and enhancing recommender systems on web platforms. The Challenge consists of two tracks: the User Modeling Track and the Recommendation Track. Participants are tasked to utilize a combined dataset from Yelp, Amazon, and Goodreads, along with an interactive environment simulator, to develop innovative LLM agents. The Challenge has attracted 295 teams across the globe and received over 1,400 submissions in total over the course of 37 official competition days. The participants have achieved 21.9% and 20.3% performance improvement for Track 1 and Track 2 in the Development Phase, and 9.1% and 15.9% in the Final Phase, representing a significant accomplishment. This paper discusses the detailed designs of the Challenge, analyzes the outcomes, and highlights the most successful LLM agent designs. To support further research and development, we have open-sourced the benchmark environment at https://tsinghua-fib-lab.github.io/AgentSocietyChallenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。