量化真实与模拟用户行为的差异,提升AI助手训练效果
Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors

- 用聚类+分布对比法测量真实与模拟用户行为差距
- 24个大模型模拟器中普遍存在显著行为偏差
- 组合互补模拟器可更接近真实用户行为
随着用户模拟器在AI助手交互式训练与评估中的广泛应用,其是否能体现真实用户行为的多样性至关重要。现有工作虽致力于生成类人响应,但真实用户行为的广泛异质性是否被准确捕捉仍存疑问。本文提出一种方法,通过从对话中提取用户行为表征,聚类生成离散分布,并计算分布差异来度量真实与模拟行为间的差距。该方法经人类研究与消融实验验证。我们首次对24个基于LLM的用户模拟器在编程与写作任务上进行系统评估,发现其与真实用户间存在显著分布差距,且差距随模型家族、规模及行为维度变化。成对比较显示多数模拟器行为相似,少数明显不同。将行为互补的模拟器结合后,其整体分布更接近真实用户。最后,基于TF-IDF的聚类分析揭示了模拟器捕获、遗漏与虚构的行为模式。
原文摘要 · Abstract (English)
As user simulators are increasingly used for interactive training and evaluation of AI assistants, it is essential that they represent the diverse behaviors of real users. While existing works train user simulators to generate human-like responses, whether they capture the broad and heterogeneous distribution of real user behaviors remains an open question. In this work, we introduce a method to measure the distributional gap between real and simulated user behaviors, validated through a human study and ablations. Given a dataset of real and simulated conversations, our method extracts representations of user behavior from each conversation, quantizes them into discrete distributions via clustering, then computes divergence metrics. We provide the first systematic evaluation of 24 LLM-based user simulators on coding and writing tasks, and reveal a large distributional gap from real users that varies across model families, scales, and behavioral facets. Pairwise comparisons show that most simulators behave similarly, while a few stand apart. Combining behaviorally complementary simulators brings the resulting distribution closer to real users compared to either simulator on its own. Finally, a TF-IDF analysis of the clusters surfaces interpretable patterns of behaviors that simulators capture, miss, and hallucinate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。