arXiv:2412.11912cs.CL2024-12

构建首个大规模双语角色定制评测基准,提升大模型角色扮演能力评估效果。

CharacterBench: Benchmarking Character Customization of Large Language Models

  • 设计11维评估体系,区分稀疏与密集特征维度,精准捕捉角色特性
  • 涵盖3956个角色、22859条人工标注样本,覆盖25类详细角色类型
  • 开发CharacterJudge模型,比GPT-4更稳定高效,适合大规模优化应用

基于角色的对话(即角色扮演)使用户能自由定制交互角色,常依赖大语言模型,因此亟需评估其角色定制能力。现有基准存在局限:通常仅涉及单一角色类别或评估维度有限,且回复中角色特征稀疏,导致以特征为核心的生成评估既无效又低效。为此,我们提出CharacterBench,目前最大的双语生成评测基准,包含22,859条人工标注样本,覆盖3,956个角色及25个详细角色类别。我们定义了11个维度,按特征是否在每条回复中显现分为稀疏与密集维度,并针对每个维度设计特定查询以诱发相关角色响应。进一步开发CharacterJudge模型,实现低成本、稳定的评估。实验表明其优于当前最优自动评判模型(如GPT-4),且该基准具备优化大模型角色定制能力的潜力。代码仓库见https://github.com/thu-coai/CharacterBench。

原文摘要 · Abstract (English)

Character-based dialogue (aka role-playing) enables users to freely customize characters for interaction, which often relies on LLMs, raising the need to evaluate LLMs' character customization capability. However, existing benchmarks fail to ensure a robust evaluation as they often only involve a single character category or evaluate limited dimensions. Moreover, the sparsity of character features in responses makes feature-focused generative evaluation both ineffective and inefficient. To address these issues, we propose CharacterBench, the largest bilingual generative benchmark, with 22,859 human-annotated samples covering 3,956 characters from 25 detailed character categories. We define 11 dimensions of 6 aspects, classified as sparse and dense dimensions based on whether character features evaluated by specific dimensions manifest in each response. We enable effective and efficient evaluation by crafting tailored queries for each dimension to induce characters' responses related to specific dimensions. Further, we develop CharacterJudge model for cost-effective and stable evaluations. Experiments show its superiority over SOTA automatic judges (e.g., GPT-4) and our benchmark's potential to optimize LLMs' character customization. Our repository is at https://github.com/thu-coai/CharacterBench.

角色扮演评测基准大模型评估生成评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。