用真人角色模拟测试大模型在零售对话中的真实客户行为表现
CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators

- 构建360个真人角色的多模态客户模拟环境,评估模型行为一致性
- 顶级模型决策对齐率不足74%,开放模型过早暴露需求
- 提出UserGRPO强化学习方法,提升决策对齐23.5点且不牺牲对话质量
我们提出CustomerSim,一个用于评估多模态大语言模型(MLLMs)在基于对话的零售环境中模拟真实、角色驱动客户行为能力的基准与环境。不同于以往仅关注表面对话生成的工作,我们聚焦模型在多轮、自主模拟中遵循客户设定进行信息搜索和决策的能力。CustomerSim包含覆盖五个产品类别的360个人工标注角色,以及衡量客户模拟行为与设定一致性的指标和对话质量评估体系。实验发现,当前五种主流开源与闭源模型存在明显行为差距:尽管对话流畅,但词汇多样性显著低于真人购物者;开源模型在首轮即过度披露标准;模型易受销售员语气影响而偏离角色设定。即使最强的闭源模型(Claude Opus 4.8 和 GPT-5.6 Sol)决策对齐率也未超过74%。为此,我们提出UserGRPO——一种多轮、多目标强化学习方案,同时优化对话流畅性与角色设定一致性。该方法将基线模型决策对齐率从0.417提升至0.652(+23.5点),且对话质量无明显下降,迁移效果良好。此外,仅风格化提示能提升表层自然度,但会几乎使角色一致性减半。通过CustomerSim,我们为社区提供了一个可探究并改进目标导向场景下用户模拟器对齐能力的测试平台。
原文摘要 · Abstract (English)
We present CustomerSim, an environment and benchmark to evaluate the extent to which Multimodal Large Language Models (MLLMs) can simulate realistic, persona-driven customer behavior in chat-based retail environments. While prior work treats user simulation as surface-level dialog generation, we focus on a model's ability to seek information and make decisions that adhere to customer specifications in multiturn, agentic simulations. CustomerSim consists of a human-curated set of 360 personas over five product categories, alongside a suite of metrics measuring consistency between a customer simulator's actions and its specifications and conversational quality. We find several behavioral gaps across five open and closed-source state-of-the-art models. First, while models produce fluent conversations, they display significantly lower lexical diversity than human shoppers, and open-source models overdisclose their criteria in the opening turn. Second, models tend to be persuaded by sales agent tone and drift from persona specifications. Even the strongest closed-source models, Claude Opus 4.8 and GPT-5.6 Sol, achieves <74% alignment with its persona specifications. To address these limitations, we propose UserGRPO, a multi-turn, multi-objective reinforcement learning recipe optimizing both conversational fluency and decision alignment under persona specifications. UserGRPO raises the decision alignment of the baseline model from 0.417 to 0.652, a gain of 23.5 points, without meaningful cost to conversational quality, and these gains transfer to held-out product categories. We further find that stylistic prompting is the only intervention that makes surface form more human-like, yet it nearly halves persona adherence. Through CustomerSim, we provide a testbed for the community to investigate and improve the adherence of user simulators in goal-oriented settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。