arXiv:2601.09215cs.CL2026-01被引 3

让用户模拟器学会人类式推理,更难被欺骗。

UserLM-R1: Modeling Human Reasoning in User Language Models with Multi-Reward Reinforcement Learning

  • 构建动态角色与目标,适应不同场景
  • 通过多奖励强化学习提升策略能力,对抗性任务表现更好
  • 适合需要真实用户交互的智能体训练场景

用户模拟器是智能体后训练的关键交互环境。理想的用户模拟器应具备跨领域泛化能力,并能主动谈判、质疑或讨价还价。但现有方法存在两大问题:依赖静态、上下文无关的用户画像,需大量人工重设新场景,泛化性差;且忽略人类策略思维,易被智能体操纵。为此,我们提出UserLM-R1,一种具备推理能力的用户语言模型。首先,构建包含静态角色与动态场景目标的综合用户画像,以适应多样化场景;其次,设计基于目标的决策策略,在生成回应前生成高质量推理过程,并通过监督微调和多奖励强化学习进一步优化推理与战略能力。大量实验表明,UserLM-R1在更具挑战性的对抗性数据集上显著优于基线模型。

原文摘要 · Abstract (English)

User simulators serve as the critical interactive environment for agent post-training, and an ideal user simulator generalizes across domains and proactively engages in negotiation by challenging or bargaining. However, current methods exhibit two issues. They rely on static and context-unaware profiles, necessitating extensive manual redesign for new scenarios, thus limiting generalizability. Moreover, they neglect human strategic thinking, leading to vulnerability to agent manipulation. To address these issues, we propose UserLM-R1, a novel user language model with reasoning capability. Specifically, we first construct comprehensive user profiles with both static roles and dynamic scenario-specific goals for adaptation to diverse scenarios. Then, we propose a goal-driven decision-making policy to generate high-quality rationales before producing responses, and further refine the reasoning and improve strategic capabilities with supervised fine-tuning and multi-reward reinforcement learning. Extensive experimental results demonstrate that UserLM-R1 outperforms competitive baselines, particularly on the more challenging adversarial set.

用户建模强化学习推理能力智能体训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。