arXiv:2502.09082cs.CLcs.AI2025-02中稿 · ICML被引 23

构建首个大规模角色扮演数据集,助力大模型真实还原文学人物言行。

CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation

  • 基于表演理论设计多角色场景训练与评估框架
  • 涵盖17,966个文学角色,支持对话、内心独白等多元数据类型
  • 开源70B模型在多个基准上超越GPT-4o,准确率达93.47%

角色扮演语言代理(RPLAs)已成为大语言模型(LLMs)的新兴应用。然而,模拟经典角色面临缺乏真实角色数据集和精细评估方法的挑战。本文提出CoSER,一个高质量数据集、开放模型及评估协议,用于有效训练与评估经典角色的角色扮演。CoSER数据集覆盖771部知名书籍中的17,966个角色,包含真实世界复杂性的对话语料,以及对话设定、角色经历与内心独白等多种数据类型。借鉴表演学方法,我们提出给定情境表演训练范式,让模型在书本场景中顺序扮演多个角色。基于此,我们开发了CoSER 8B与CoSER 70B两个基于LLaMA-3的先进开源角色扮演模型。大量实验表明,CoSER数据集在角色扮演训练、评估与检索中具有显著价值。此外,CoSER 70B在我们的评估及三个现有基准上表现领先,于InCharacter与LifeChoice基准分别取得75.80%与93.47%的准确率。

原文摘要 · Abstract (English)

Role-playing language agents (RPLAs) have emerged as promising applications of large language models (LLMs). However, simulating established characters presents a challenging task for RPLAs, due to the lack of authentic character datasets and nuanced evaluation methods using such data. In this paper, we present CoSER, a collection of a high-quality dataset, open models, and an evaluation protocol towards effective RPLAs of established characters. The CoSER dataset covers 17,966 characters from 771 renowned books. It provides authentic dialogues with real-world intricacies, as well as diverse data types such as conversation setups, character experiences and internal thoughts. Drawing from acting methodology, we introduce given-circumstance acting for training and evaluating role-playing LLMs, where LLMs sequentially portray multiple characters in book scenes. Using our dataset, we develop CoSER 8B and CoSER 70B, i.e., advanced open role-playing LLMs built on LLaMA-3.1 models. Extensive experiments demonstrate the value of the CoSER dataset for RPLA training, evaluation and retrieval. Moreover, CoSER 70B exhibits state-of-the-art performance surpassing or matching GPT-4o on our evaluation and three existing benchmarks, i.e., achieving 75.80% and 93.47% accuracy on the InCharacter and LifeChoice benchmarks respectively.

角色扮演大模型文本生成数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。