arXiv:2505.13157cs.CLcs.AI2025-05被引 4

构建新评测基准,全面评估大模型角色扮演能力。

Role-Playing Evaluation for Large Language Models

  • 设计四维评测体系:情绪理解、决策能力、道德一致性、角色一致
  • 提供可复现的基准数据集与评估代码,支持自动化测试
  • 适合研究角色扮演、对齐性及生成质量的AI开发者使用

大语言模型在扮演角色方面展现出显著能力,但评估这一能力面临挑战:人工评价成本高,自动化评估易有偏差。为此,我们提出角色扮演评测(RPEval),一个针对大模型角色扮演能力的新基准,涵盖情感理解、决策能力、道德对齐和角色一致性四个关键维度。本文详细介绍了RPEval的构建过程,并给出了基线评估结果。相关代码与数据集已公开于https://github.com/yelboudouri/RPEval。

原文摘要 · Abstract (English)

Large Language Models (LLMs) demonstrate a notable capacity for adopting personas and engaging in role-playing. However, evaluating this ability presents significant challenges, as human assessments are resource-intensive and automated evaluations can be biased. To address this, we introduce Role-Playing Eval (RPEval), a novel benchmark designed to assess LLM role-playing capabilities across four key dimensions: emotional understanding, decision-making, moral alignment, and in-character consistency. This article details the construction of RPEval and presents baseline evaluations. Our code and dataset are available at https://github.com/yelboudouri/RPEval

角色扮演评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。