arXiv:2605.02624cs.CL2026-05被引 1

提出新框架评估对话模拟器真实度,发现模拟用户常忽略真实摩擦。

Synthetic Users, Real Differences: an Evaluation Framework for User Simulation in Multi-Turn Conversations

论文配图:Synthetic Users, Real Differences: an Evaluation Framework for User Simulation in Multi-Turn Conversations
图 1 · 摘自论文原文
  • 从8个维度对比真实与模拟对话的分布差异
  • 1000组多轮任务对话显示模拟对话普遍高估系统表现
  • 不同领域表现差异大,需定制化模拟器

用户模拟在聊天机器人评估中日益受到关注,可替代真实用户交互数据。为提升评估严谨性,本文提出realsim框架,支持从业务功能、用户状态和话语形式等8个维度对真实与模拟对话进行分布级比较。基于涵盖16个应用领域的1000组多轮任务对话数据集,研究发现模拟用户难以捕捉真实用户带来的沟通摩擦,导致基于模拟的评估结果往往过于乐观。同时,各领域间表现差异显著,提示需构建领域专用的用户模拟器。

原文摘要 · Abstract (English)

There is growing interest in exploring user simulation as an alternative to gathering and scoring real user-chatbot interactions for AI chatbot evaluation. For this purpose, it is important to ensure the realism of the simulation, i.e., the extent to which simulated dialogues reflect real dialogues users have with chatbots. Most existing methods evaluating simulation realism produce coarse quality signal and remain solely at the level of individual dialogues. To support more rigorous evaluation in this area, we propose realsim, an evaluation framework that enables practitioners to take a distributional view of real vs. simulated dialogues along 8 dimensions, covering attributes related to the communicative functions of the interaction, user states, and the surface form of user messages. We then instantiate the framework with a curated dataset of 1K multi-turn task-focused real user-chatbot dialogues that cover 16 domains of chatbot applications. Overall, we find that simulated users tend to struggle at capturing communication frictions that real users introduce to interactions, which could make evaluations based on such simulations overly optimistic. We also observe variability in performance across different domains, which may indicate a need for domain-specific user simulators.

用户模拟对话评估多轮对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。