发现大模型用户模拟器与真实人类行为存在显著差距,易导致智能体评估虚高。
Mind the Sim2Real Gap in User Simulation for Agentic Tasks
- 首次系统验证用户模拟中的仿真到现实差距,引入用户模拟指数(USI)量化偏差。
- 31个大模型模拟器普遍过度合作、反馈过于正面,使智能体成功率虚高。
- 适合智能体研发者关注,提醒在评估中必须用真人验证模拟器可靠性。
随着自然语言处理评估从静态基准转向多轮交互场景,基于大模型的用户模拟器被广泛用作用户代理,承担生成用户对话和提供评估信号双重角色。然而,这些模拟常被默认忠实于真实人类行为,缺乏严格验证。本文首次形式化了用户模拟中的仿真到现实差距,并通过真实人类实验(451名参与者,165项任务)完整执行τ-基准协议,对31个来自专有、开源及专用家族的大模型模拟器进行评估。我们提出用户模拟指数(USI),用于量化大模型模拟器与真实用户交互行为及反馈的相似度。结果表明,大模型模拟器行为上过于合作、风格统一,缺乏真实人类的挫败感与模糊性,形成‘简单模式’,使智能体成功率达人类基线之上。在评估中,真人用户在八个质量维度给出细致判断,而模拟用户则产生高度一致的正面反馈;基于规则的奖励机制无法捕捉人类用户产生的丰富反馈信号。总体而言,模型通用能力越强,其用户模拟真实性未必越高。研究强调在智能体开发流程中使用大模型用户模拟器时,必须进行人类验证,并推动更精准的用户模拟模型发展。
原文摘要 · Abstract (English)
As NLP evaluation shifts from static benchmarks to multi-turn interactive settings, LLM-based simulators have become widely used as user proxies, serving two roles: generating user turns and providing evaluation signals. Yet, these simulations are frequently assumed to be faithful to real human behaviors, often without rigorous verification. We formalize the Sim2Real gap in user simulation and present the first study running the full $τ$-bench protocol with real humans (451 participants, 165 tasks), benchmarking 31 LLM simulators across proprietary, open-source, and specialized families using the User-Sim Index (USI), a metric we introduce to quantify how well LLM simulators resemble real user interactive behaviors and feedback. Behaviorally, LLM simulators are excessively cooperative, stylistically uniform, and lack realistic frustration or ambiguity, creating an "easy mode" that inflates agent success rates above the human baseline. In evaluations, real humans provide nuanced judgments across eight quality dimensions while simulated users produce uniformly more positive feedback; rule-based rewards are failing to capture rich feedback signals generated by human users. Overall, higher general model capability does not necessarily yield more faithful user simulation. These findings highlight the importance of human validation when using LLM-based user simulators in the agent development cycle and motivate improved models for user simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。