arXiv:2605.20204cs.HCcs.AI2026-05被引 4

用真实人类对话数据构建用户模拟器,让智能体评测更贴近现实。

RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation

论文配图:RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
图 1 · 摘自论文原文
  • 基于1.4万条真实对话提取7275个可执行行为模板,驱动模拟器
  • 行为匹配率从24.2%提升至45.3%,在71个领域验证有效
  • 揭露传统模拟器的夸大指令问题,适合评估真实场景下的智能体

基于大模型的用户模拟是智能体端到端评估的核心机制,但现有模拟器难以代表真实人类:无约束的LLM默认行为导致形式匹配率仅6-8%;手工指令则引发指令放大效应,使模型过度解读生成不自然的行为极端。本文提出RealUserSim,首个基于真实行为数据的用户模拟框架。从超过1.4万条真实人机对话(WildChat)中提取7,275个可执行行为配置,并用于训练模拟器。在包含600条对话、71+领域的反泄露控制基准(PT3)测试中,真实数据驱动的模拟将五维行为匹配率从24.2%提升至45.3%。在TauBench上的多模型评估表明,该框架能作为真实压力测试,暴露合作型模拟器无法发现的三种失败模式(任务成功率下降3.2%至3.5%),而现有基准中的指令放大效应导致行为失真,损害评估有效性。

原文摘要 · Abstract (English)

LLM-based user simulation is the primary mechanism for end-to-end agent evaluation, yet simulated users are poor proxies for real humans: unconstrained LLM defaults produce a Formalism Ceiling (style match rates of 6-8% against real users), while hand-crafted behavioral directives trigger Directive Amplification, where models hyper-interpret instructions into unnatural behavioral extremes that vary dramatically across simulator models. We present RealUserSim, the first user simulation framework grounded in real behavioral data. From 14,000+ authentic human-LLM conversations (WildChat), we extract 7,275 executable behavioral profiles and use them to ground LLM simulators. A fidelity benchmark (PT3) on 600 conversations across 71+ domains with anti-leakage controls shows that grounded simulation raises match rate from 24.2% to 45.3% across five behavioral dimensions. Agent evaluation on TauBench with 6 simulator models and extensive analysis shows that grounded simulation acts as a realistic stress test, surfacing three failure mechanisms invisible to cooperative simulators (mean -3.2% to -3.5% task success degradation), while Directive Amplification in existing benchmarks produces unrealistic behavior that compromises the validity of agent evaluation.

用户模拟智能体评测真实数据行为匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。