用多模型模拟用户,评测大模型角色扮演能力。
PingPong: A Benchmark for Role-Playing Language Models with User Emulation and Multi-Model Evaluation
- 用不同模型模拟用户行为,动态测试角色扮演表现。
- 40+模型参与64轮对话,3项指标验证效果一致。
- 适合评估交互式场景下的模型真实表现能力。
我们提出一个用于评估语言模型角色扮演能力的基准测试。方法包括:扮演特定角色的玩家模型、模拟用户行为的情境化提问模型,以及由多个评分模型组成的评判系统,从角色一致性、娱乐性和语言流畅性三个维度评估对话质量。我们对超过40个模型在英语和俄语中进行了评估,每个模型参与64次对话,涵盖8种角色和8种情境。通过对比自动化评估与人工标注结果,验证了该方法在多个指标上的强相关性。本工作为交互场景下模型能力的稳健、动态评估提供了基础。
原文摘要 · Abstract (English)
We introduce a benchmark for evaluating the role-playing capabilities of language models. Our approach leverages different language models to simulate users in dynamic, multi-turn conversations and assess the resulting dialogues. Our methodology involves three main components: a player model that adopts a specific character role, an interrogator model that simulates user behavior in a specific situation, and a judge model ensemble that evaluates conversation quality with 3 metrics: character consistency, entertainment value, and language fluency. We evaluated more than 40 models in both English and Russian, with each model participating in 64 conversations with 8 characters and 8 situations. We conducted experiments comparing automated evaluations with human annotations to validate our approach, demonstrating strong correlations across multiple criteria. This work provides a foundation for a robust and dynamic evaluation of different model capabilities in interactive scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。