测试大模型模拟用户决策的真实度,发现其会虚假提升拒绝用户的购买意愿。
Simulated Customers Never Walk Away: Decision Fidelity of LLM User Simulators Measured Against Real Purchase Outcomes
- 提出'决策真实度'新标准,评估模拟用户是否像真人一样在真实选择中逐步放弃
- 实测2790条对话发现,模拟用户对不买者的阻力降低50%,拖延时间翻倍却无真实付款
- 适合训练销售智能体的研究者,避免因模拟偏差而高估转化效果
LLM作为用户模拟器已成对话AI核心基础设施:代理基准(tau-bench)、训练流程及大量真实性研究均依赖大模型扮演对话中的人类角色。现有框架仅衡量沟通真实性——模拟器是否像真人说话——基于付费参与者扮演预设目标的真值数据。我们指出此方法存在结构性盲点:当目标被分配时,用户意愿为外生变量,无法检验模拟器能否复制真实用户内生、隐性且随时间衰减的决策动态。为此引入决策真实度——模拟群体是否再现真实用户面对真实重要选择时的决策状态演化。在包含2,790条真实客户与LLM销售代理的生产对话数据集上,其中793例有验证付款结果,采用教师强制探测协议,在固定上下文和工具条件下,发现系统性、结果相关的失败现象,称作‘脱钩缺陷’:模拟器几乎准确复现最终买家(深度偏差+0.09),但将最终未购买者推向购买框架(深度偏差+0.40;d=0.38,p<0.001),使表达抗拒率从25.1%降至13.5%(下降50%),拖延比例从21.9%升至40.1%(翻倍),却虚构出零付款。该缺陷在不同模型族(DeepSeek:d=0.41,p=0.002)中重复出现,且即使指令模拟器可中途退出,边际偏差仅降低五倍,但结果条件差异仍显著(d=0.34,p=0.008)。真实用户说‘不现在’就停止;模拟用户则追问价格。用此类模拟器评估或训练销售与说服型智能体,将在最关键的流失客户环节错误夸大转化进展。
原文摘要 · Abstract (English)
LLM-as-user-simulation has become core infrastructure for conversational AI: agent benchmarks (tau-bench), training pipelines, and a growing body of fidelity studies all rely on LLMs role-playing the human side of dialogue. Existing frameworks measure communicative fidelity -- whether simulators talk like humans -- against ground truth from paid participants role-playing assigned goals. We argue this has a structural blind spot: when the goal is assigned, the user's willingness is exogenous, so no framework can test whether simulators make decisions like real users whose motivation is endogenous, latent, and decaying. We introduce decision fidelity -- whether a simulated population reproduces the decision-state dynamics of real users facing real, consequential choices -- and measure it on a unique testbed: 2,790 production conversations between an LLM sales agent and real customers, including 793 with verified payment outcomes. Using a teacher-forced probe protocol that holds context and instrument fixed, we find a systematic, outcome-correlated failure we call the disengagement deficit: simulators reproduce eventual buyers almost exactly (depth bias +0.09) but inflate eventual non-buyers toward the purchase frame (depth bias +0.40; d=0.38, p<0.001), halving expressed resistance (25.1% to 13.5%) and nearly doubling deliberation (21.9% to 40.1%) while fabricating no purchases. The deficit replicates across model families (DeepSeek: d=0.41, p=0.002) and resists the obvious fix: instructing the simulator that it may disengage cuts marginal bias five-fold but barely moves the outcome-conditioned contrast (d=0.34, p=0.008). Real non-buyers say "not now" and stop; simulated non-buyers ask about price. Evaluating or training sales and persuasion agents against such simulators overstates funnel progress exactly where it matters most -- the customers who walk away.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。