arXiv:2510.06552cs.CL2025-10中稿 · ICLR被引 46

用专门训练的用户模型更真实地模拟对话,让助手表现更差才说明评估更可信。

Flipping the Dialogue: Training and Evaluating User Language Models

  • 训练专用用户语言模型,替代助手模型模拟人类对话者。
  • 真实对话模拟下,强助手性能从74.6%降至57.4%。
  • 适合评估对话系统在复杂多轮场景下的真实鲁棒性。

与大模型进行对话涉及两个参与者:人类用户主导对话,大模型助手响应用户请求。为胜任此角色,大模型经过后训练,优化为能提供详尽、结构清晰、无歧义和语法错误的回应。而用户话语则很少完美,每个人表达方式不同,常在对话中逐步调整措辞。以往研究通过让原本作为助手训练的大模型扮演用户来模拟多轮对话。然而我们发现,越优秀的助手模型,其作为用户模拟器的表现越差。为此,我们提出专为模拟人类用户设计的用户语言模型(User LMs),通过多种评估验证其行为更贴近真实用户,模拟鲁棒性更强。在编码和数学对话场景中,使用用户模型模拟后,强助手GPT-4o性能从74.6%下降至57.4%,表明更真实的模拟环境更能暴露助手在多轮对话中的局限。

原文摘要 · Abstract (English)

Conversations with LMs involve two participants: a human user leading the conversation, and an LM assistant responding to the user's request. To satisfy this specific role, LMs are post-trained to be helpful assistants -- optimized to produce exhaustive and well-structured responses, free of ambiguity and grammar errors. User utterances, on the other hand, are rarely perfected, with each user phrasing requests in unique ways, sometimes putting in partial effort at each turn and refining on the fly. To evaluate LM performance in realistic settings, prior work simulated users in multi-turn conversations, often by prompting an LM originally trained to be a helpful assistant to act as a user. However, we show that assistant LMs make for poor user simulators, with the surprising finding that better assistants yield worse simulators. Instead, we introduce purpose-built User Language Models (User LMs) - models post-trained to simulate human users in multi-turn conversations. Through various evaluations, we show how User LMs align better with human behavior and achieve better simulation robustness than existing simulation methods. When leveraging User LMs to simulate coding and math conversations, the performance of a strong assistant (GPT-4o) drops from 74.6% to 57.4%, confirming that more realistic simulation environments lead to assistant struggles as they fail to cope with the nuances of users in multi-turn setups.

对话评估用户模拟大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。