测试用户模拟器能否可靠替代真人评估聊天机器人表现。
SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?
- 用真实对话数据构建模拟器评测基准,对比模拟与真人行为匹配度。
- 基于用户画像的模拟器在双任务中相关性达斯皮尔曼ρ=0.7,接近真人评分。
- 适用于快速评估多轮交互模型,尤其适合缺乏真人评测资源的研究者。
大语言模型(LLM)在交互式应用中日益普及,而人类评估仍是多轮对话性能的黄金标准。由于人力评估成本高、耗时长且难以复现,近期研究尝试使用LLM模拟用户以实现自动评估。然而,目前尚无基准或系统性研究来检验这些模拟用户是否能可靠替代真实用户。为此,我们提出SimulatorArena,一个包含909条标注的人类-LLM对话的数据集,涵盖数学辅导和文档创作两个交互任务。该基准通过衡量模拟消息与人类行为的相似度,以及模拟评分与人类判断的一致性来评估模拟器性能。实验表明,基于用户画像(如背景、表达风格)训练的模拟器在两项任务中均达到斯皮尔曼等级相关系数ρ=0.7,提供了可扩展、实用的替代方案。利用最佳模拟器,我们对18个助手(包括GPT-5、Claude 4.1 Opus、Gemini 2.5 Pro等最新模型)进行了评测。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in interactive applications, and human evaluation remains the gold standard for assessing their performance in multi-turn conversations. Since human studies are costly, time-consuming, and hard to reproduce, recent work explores using LLMs to simulate users for automatic assistant evaluation. However, there is no benchmark or systematic study to evaluate whether these simulated users are reliable stand-ins for real users. To address this, we introduce SimulatorArena, a benchmark of 909 annotated human-LLM conversations on two interactive tasks -- math tutoring and document creation. SimulatorArena evaluates simulators based on how closely their messages match human behavior and how well their assistant ratings align with human judgments. Experiments on various simulator methods show that simulators conditioned on user profiles, capturing traits like background and message styles, align closely with human judgments. They reach Spearman's $ρ$ of 0.7 on both tasks, providing a practical, scalable alternative to human evaluation. Using the best simulator for each task, we benchmark 18 assistants, including the latest LLMs such as GPT-5, Claude 4.1 Opus, and Gemini 2.5 Pro.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。