用图灵测试思想训练更像真人的用户模拟器。
Learning User Simulators with Turing Rewards

- 用大模型做裁判,评估生成回复与真人是否难分真假。
- 在聊天和Reddit讨论中,效果优于传统匹配方法。
- 适合需要逼真交互的智能助手训练与评测场景。
在交互场景中学习模拟人类用户,可推动代理助手训练、个性化系统评估及社会科学研究。现有方法通常通过训练大语言模型(LLM)以匹配单一真实回复,或最大化对数概率,或使用相似性奖励。我们提出基于图灵测试的强化学习方法Turing-RL:利用大模型裁判生成判别式图灵奖励,评估生成回复在用户历史背景下的真实性,使用户模拟器学习生成难以与真实用户回复区分的内容。在对话聊天和Reddit论坛讨论两个领域,Turing-RL在大模型和人工评估指标上均持续优于基线方法。研究表明,优化不可区分性而非单纯匹配,是学习用户模拟器的有效策略。
原文摘要 · Abstract (English)
Learning to simulate human users in interactive settings could advance the training of agent assistants, evaluation of personalization systems, research in the social sciences, and more. Existing approaches generally do so by training a large language model (LLM) to match a single ground truth response, either by maximizing the log probability or by using a similarity reward. We instead propose Turing-RL: a Turing-Test-based reinforcement learning approach for training user simulator models. Turing-RL uses a discriminative Turing reward with an LLM judge to score how indistinguishable a generated response is from the real user's given the user's history, and the user simulator LLM learns to produce responses indistinguishable from what the user could have said with such rewards. Across two different domains--conversational chat and Reddit forum discussion--we find that Turing-RL consistently outperforms baseline methods on both LLM and human evaluation metrics. Our study suggests that optimizing for indistinguishability, rather than response matching, is effective for learning user simulators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。