arXiv:2510.07230cs.CL2025-10被引 19

用强化学习让大模型更真实地模拟用户的个性化购物行为。

Customer-R1: Personalized Simulation of Human Behaviors via RL-based LLM Agent in Online Shopping

  • 基于强化学习,让大模型根据用户画像生成下一步行为。
  • 在OPeRA数据集上,预测准确率显著优于提示和微调方法。
  • 适合研究个性化推荐与用户行为建模的从业者。

利用大语言模型(LLMs)模拟逐步的人类行为已成为新兴研究方向,广泛应用于各类实际场景。尽管提示、监督微调(SFT)和强化学习(RL)等方法在建模步骤化行为方面已展现潜力,但它们主要学习的是群体层面的策略,未考虑用户个性特征,导致模拟结果泛化而非个性化。本文提出核心问题:如何让LLM代理更好地模拟个性化用户行为?为此,我们引入Customer-R1,一种基于强化学习的在线购物环境中个性化、步骤化用户行为模拟方法。其策略显式依赖用户画像,并通过动作正确性奖励信号优化下一步推理与动作生成。在OPeRA数据集上的实验表明,Customer-R1不仅在下一步动作预测任务中显著优于提示和SFT基线,且更贴近用户的动作分布,显示出更高的个性化行为模拟保真度。

原文摘要 · Abstract (English)

Simulating step-wise human behavior with Large Language Models (LLMs) has become an emerging research direction, enabling applications in various practical domains. While prior methods, including prompting, supervised fine-tuning (SFT), and reinforcement learning (RL), have shown promise in modeling step-wise behavior, they primarily learn a population-level policy without conditioning on a user's persona, yielding generic rather than personalized simulations. In this work, we pose a critical question: how can LLM agents better simulate personalized user behavior? We introduce Customer-R1, an RL-based method for personalized, step-wise user behavior simulation in online shopping environments. Our policy is conditioned on an explicit persona, and we optimize next-step rationale and action generation via action correctness reward signals. Experiments on the OPeRA dataset emonstrate that Customer-R1 not only significantly outperforms prompting and SFT-based baselines in next-action prediction tasks, but also better matches users' action distribution, indicating higher fidelity in personalized behavior simulation.

个性化建模强化学习行为模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。