构建新评测基准,全面评估大模型角色扮演能力。
Role-Playing Evaluation for Large Language Models
- 设计四维评测体系:情绪理解、决策能力、道德一致性、角色一致
- 提供可复现的基准数据集与评估代码,支持自动化测试
- 适合研究角色扮演、对齐性及生成质量的AI开发者使用
大语言模型在扮演角色方面展现出显著能力,但评估这一能力面临挑战:人工评价成本高,自动化评估易有偏差。为此,我们提出角色扮演评测(RPEval),一个针对大模型角色扮演能力的新基准,涵盖情感理解、决策能力、道德对齐和角色一致性四个关键维度。本文详细介绍了RPEval的构建过程,并给出了基线评估结果。相关代码与数据集已公开于https://github.com/yelboudouri/RPEval。
原文摘要 · Abstract (English)
Large Language Models (LLMs) demonstrate a notable capacity for adopting personas and engaging in role-playing. However, evaluating this ability presents significant challenges, as human assessments are resource-intensive and automated evaluations can be biased. To address this, we introduce Role-Playing Eval (RPEval), a novel benchmark designed to assess LLM role-playing capabilities across four key dimensions: emotional understanding, decision-making, moral alignment, and in-character consistency. This article details the construction of RPEval and presents baseline evaluations. Our code and dataset are available at https://github.com/yelboudouri/RPEval
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。