arXiv:2604.09549cs.IRcs.AI2026-04被引 3

用带生活场景的智能体模拟真实用户,提升推荐系统评估准确性

Beyond Offline A/B Testing: Context-Aware Agent Simulation for Recommender System Evaluation

论文配图:Beyond Offline A/B Testing: Context-Aware Agent Simulation for Recommender System Evaluation
图 1 · 摘自论文原文
  • 构建基于日常生活场景的智能体,模拟时间、地点和需求等上下文
  • 生成的行为数据与真实人类更接近,优化后可提升实际点击率
  • 适合研究推荐系统评估、个性化建模或想替代传统离线测试的研究者

推荐系统在在线服务中至关重要,帮助用户从海量内容中导航。然而,其评估仍面临离线指标与线上表现脱节的问题。大语言模型驱动的智能体提供了新思路,但现有研究将用户行为孤立建模,忽视了时间、位置和需求等关键上下文因素。本文提出 ContextSim 框架,通过生活情景模块生成何时、何地、为何使用推荐的具体场景,使智能体代理具备真实用户行为特征。为确保偏好一致性,我们在动作和行为轨迹层面建模内部思考并施加约束。跨领域实验表明,该方法生成的交互更贴近人类行为;通过离线 A/B 测试相关性验证,基于 ContextSim 优化的推荐系统参数能显著提升真实世界用户参与度。

原文摘要 · Abstract (English)

Recommender systems are central to online services, enabling users to navigate through massive amounts of content across various domains. However, their evaluation remains challenging due to the disconnect between offline metrics and online performance. The emergence of Large Language Model-powered agents offers a promising solution, yet existing studies model users in isolation, neglecting the contextual factors such as time, location, and needs, which fundamentally shape human decision-making. In this paper, we introduce ContextSim, an LLM agent framework that simulates believable user proxies by anchoring interactions in daily life activities. Namely, a life simulation module generates scenarios specifying when, where, and why users engage with recommendations. To align preferences with genuine humans, we model agents' internal thoughts and enforce consistency at both the action and trajectory levels. Experiments across domains show our method generates interactions more closely aligned with human behavior than prior work. We further validate our approach through offline A/B testing correlation and show that RS parameters optimized using ContextSim yield improved real-world engagement.

推荐系统智能体仿真评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。