arXiv:2507.20352cs.CL2025-07EMNLP被引 9

新基准RMTBench从用户意图出发,评测大模型角色扮演能力。

RMTBench: Benchmarking LLMs Through Multi-Turn User-Centric Role-Playing

  • 以用户动机为核心设计对话,而非仅看角色设定
  • 涵盖80个角色、超8000轮对话,支持多轮真实场景模拟
  • 适合评估实际应用中角色扮演的自然性与实用性

大语言模型在角色扮演应用中展现巨大潜力,但评估仍具挑战。现有基准多采用角色中心视角,将用户-角色互动简化为孤立问答,难以反映真实场景。为此,我们提出RMTBench,一个面向用户的双语角色扮演基准,包含80个多样化角色和超过8,000轮对话。角色涵盖详细背景的定制角色与仅由简单特质定义的抽象角色,支持多种用户场景评估。对话基于明确用户动机构建,而非角色描述,确保与实际应用对齐。我们设计了真实的多轮对话仿真机制,结合精心选取的评估维度与基于LLM的评分系统,有效捕捉用户与角色间复杂意图交互。通过将评估重点从角色背景转向用户意图实现,RMTBench弥合了学术评测与实际部署之间的差距,为评估大模型角色扮演能力提供了更有效的框架。代码与数据集将公开,地址:https://huggingface.co/datasets/xiangh/RMTBENCH。

原文摘要 · Abstract (English)

Recent advancements in Large Language Models (LLMs) have shown outstanding potential for role-playing applications. Evaluating these capabilities is becoming crucial yet remains challenging. Existing benchmarks mostly adopt a \textbf{character-centric} approach, simplify user-character interactions to isolated Q&A tasks, and fail to reflect real-world applications. To address this limitation, we introduce RMTBench, a comprehensive \textbf{user-centric} bilingual role-playing benchmark featuring 80 diverse characters and over 8,000 dialogue rounds. RMTBench includes custom characters with detailed backgrounds and abstract characters defined by simple traits, enabling evaluation across various user scenarios. Our benchmark constructs dialogues based on explicit user motivations rather than character descriptions, ensuring alignment with practical user applications. Furthermore, we construct an authentic multi-turn dialogue simulation mechanism. With carefully selected evaluation dimensions and LLM-based scoring, this mechanism captures the complex intention of conversations between the user and the character. By shifting focus from character background to user intention fulfillment, RMTBench bridges the gap between academic evaluation and practical deployment requirements, offering a more effective framework for assessing role-playing capabilities in LLMs. All code and datasets will be released soon. We release the datasets at https://huggingface.co/datasets/xiangh/RMTBENCH.

角色扮演大模型评测多轮对话用户意图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。