用真实用户对话训练模拟器,能显著提升聊天助手的实战表现。
Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants

- 用真实人类对话数据微调模拟器,替代单纯角色扮演的LLM。
- 微调模拟器训练出的助手在真人测试中胜率57%,远超角色扮演模拟器的51%。
- 模拟器质量应以对真实用户的效果衡量,而非仅看生成逼真度。
用户模拟器被广泛用于构建交互式AI助手,但其质量如何评估仍不明确。本文通过控制实验,考察不同模拟器对下游助手性能的影响:从提示角色扮演的LLM到在WildChat真实对话数据上微调的模拟器。在包含283名参与者的用户研究和基于真实人机对话的WildBench基准上评估,基于角色扮演的模拟器训练出的助手与初始助手无显著差异(胜率51%),而基于微调模拟器训练的助手胜率高达58%(相比初始)和57%(相比角色扮演)。进一步分析发现:角色扮演的改进策略(如人格条件化)无法弥补差距;扩大模型规模仅对微调模拟器有益;角色扮演训练的助手无法跨模拟器泛化,而微调模拟器训练的则可。结果表明,用户模拟器应基于真实行为,并以对真实用户的实际效果来衡量质量。
原文摘要 · Abstract (English)
User simulators are increasingly leveraged to build interactive AI assistants, yet how to measure the quality of these simulators remains an open question. In this work, we show how simulator quality can be quantified in terms of its downstream utility: how an LLM assistant trained with this user simulator performs in the wild when interacting with real humans. In a controlled experiment where only the user simulator varies, we train LLM assistants via reinforcement learning against a spectrum of simulators, from an LLM prompted to role-play a user to one fine-tuned on human utterances from WildChat. As evaluation, we measure pairwise win rates in a user study with 283 participants and on WildBench, a benchmark derived from real human--AI conversations. Training against the role-playing LLM yields an assistant statistically indistinguishable from the initial assistant in our user study (51% win rate), whereas training against the fine-tuned simulator yields significant gains (58% over the initial and 57% over the one trained against role-playing). Closer inspection reveals three further patterns: methods for making role-playing LLMs more realistic (e.g., persona conditioning) improve trained assistants but do not close the gap to the fine-tuned simulator; scaling the simulator's model size benefits the fine-tuned simulator but yields no gain for role-playing ones; and assistants trained against role-playing simulators fail to generalize when paired with other simulators at test time, while the one trained against fine-tuned simulator does. Together, these results argue for grounding user simulators in real human behavior and measuring their quality by their downstream effect on real users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。