arXiv:2603.03303cs.CLcs.AI2026-03被引 23

用心理状态对齐训练大模型,让虚拟用户更像真人。

HumanLM: Simulating Users with State Alignment Beats Response Imitation

  • 通过强化学习让模型生成与真实反应匹配的心理状态描述
  • 在6个数据集上平均提升16.3%的对齐得分,优于现有方法
  • 适合需要真实用户行为模拟的交互系统研究者

大型语言模型(LLMs)被用于模拟特定用户在给定情境下的回应,以支持更以用户为中心的应用。然而,现有用户模拟器多仅模仿表层语言风格,未能反映真实用户的内在状态(如信念和情绪)。为此,我们提出HumanLM框架,构建能准确反映真实用户行为的模拟器。核心思想是:除生成回应外,模型还需通过强化学习生成与真实回应对齐的自然语言潜态,这些潜态对应一组心理学基础的状态维度,驱动真实用户的行为。HumanLM进一步将对齐的潜态合成回应,实现精准模拟。为全面评估,我们构建了Humanual基准,基于公开数据涵盖6个大规模数据集,共26,000名用户、216,000条回应,覆盖日常生活、政治博客及与大模型助手的对话等任务。在各数据集上,HumanLM显著优于其他方法,平均相对对齐分提升16.3%(由大模型裁判评估)。在111名参与者的真实场景模拟测试中,HumanLM生成回应最接近真实用户,且人类相似度评分具有竞争力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used to simulate how specific users respond to a given context, enabling more user-centric applications that rely on user feedback. However, existing user simulators mostly imitate surface-level patterns and language styles, which fail to reflect the underlying states of real users (e.g., beliefs and emotions). To address these limitations, we propose a novel training framework, HumanLM, which builds user simulators that accurately reflect real users. Our key insight is that, in addition to generating responses, the model should generate natural-language latent states that align with ground-truth responses through reinforcement learning. These latent states correspond to a set of psychologically grounded state dimensions that drive how real users respond. HumanLM further synthesizes these aligned latent states into responses that accurately represent real users. For extensive evaluation, we develop Humanual, a comprehensive benchmark for simulating real users based on public data. Humanual consists of six large-scale datasets with 26k users and 216k responses in total, spanning diverse tasks such as generating user responses to daily life issues, political blogs, and chat sessions with LLM assistants. Across datasets, HumanLM significantly outperforms alternative approaches, achieving an average relative improvement of 16.3% in alignment scores from an LLM judge. In a real-time simulation study with 111 participants, HumanLM achieves the highest similarity to real user responses and competitive human-likeness scores.

用户模拟心理建模强化学习大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。