用世界模型训练更像真人的推荐系统评估代理。
AlignUSER: Human-Aligned LLM Agents via World Models for Recommender System Evaluation
- 基于人类交互数据学习世界模型,提升代理对环境理解。
- 在多个数据集上,代理行为与真实用户更接近。
- 适合需要高保真人效评估的推荐系统研究者。
推荐系统评估仍面临离线指标与真实用户行为之间的差距,以及交互数据稀缺的问题。现有工作尝试用大语言模型(LLM)代理作为合成用户,但通常依赖少样本提示,对环境理解浅显,难以忠实复现用户行为。我们提出AlignUSER框架,通过人类交互的行动与状态序列,学习由世界模型驱动的代理。将世界建模形式化为下一状态预测任务,帮助代理内化环境。为对齐行为与人类人格,围绕示范生成反事实轨迹,引导LLM比较自身决策与人类选择,识别次优动作并提取经验。所学策略用于驱动代理与推荐系统交互。我们在多个数据集上评估AlignUSER,结果表明其在微观和宏观层面均比之前方法更贴近真实人类。
原文摘要 · Abstract (English)
Evaluating recommender systems remains challenging due to the gap between offline metrics and real user behavior, as well as the scarcity of interaction data. Recent work explores large language model (LLM) agents as synthetic users, yet they typically rely on few-shot prompting, which yields a shallow understanding of the environment and limits their ability to faithfully reproduce user actions. We introduce AlignUSER, a framework that learns world-model-driven agents from human interactions. Given rollout sequences of actions and states, we formalize world modeling as a next state prediction task that helps the agent internalize the environment. To align actions with human personas, we generate counterfactual trajectories around demonstrations and prompt the LLM to compare its decisions with human choices, identify suboptimal actions, and extract lessons. The learned policy is then used to drive agent interactions with the recommender system. We evaluate AlignUSER across multiple datasets and demonstrate closer alignment with genuine humans than prior work, both at the micro and macro levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。