arXiv:2501.15427cs.CL2025-01被引 29

用合成角色数据训练可定制的对话大模型,效果媲美GPT-4o。

OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas

  • 基于人物库生成大规模角色画像,用重写与生成策略构建对齐角色的指令数据。
  • 在LLaMA-3 8B上微调后,角色扮演表现接近GPT-4o水平。
  • 开源合成角色与对话数据,助力公开研究与应用开发。

可定制的角色扮演(即角色泛化)因其在对话智能体开发与部署中的灵活性与低成本,正受到越来越多关注。本文提出一种大规模数据合成方法,使大语言模型具备角色泛化能力。我们首先利用Persona Hub中的人物资料生成大规模角色画像,随后采用响应重写与响应生成两种策略,构建与角色对齐的指令式回复。为验证合成指令数据在角色泛化上的有效性,我们基于LLaMA-3 8B模型进行监督微调(SFT)。最佳模型显著提升原始LLaMA-3 8B Instruct性能,角色对话表现达到与GPT-4o相当的水平。相关合成角色与指令对话数据已开源,以支持公共研究。

原文摘要 · Abstract (English)

Customizable role-playing in large language models (LLMs), also known as character generalization, is gaining increasing attention for its versatility and cost-efficiency in developing and deploying role-playing dialogue agents. This study explores a large-scale data synthesis approach to equip LLMs with character generalization capabilities. We begin by synthesizing large-scale character profiles using personas from Persona Hub and then explore two strategies: response rewriting and response generation, to create character-aligned instructional responses. To validate the effectiveness of our synthetic instruction tuning data for character generalization, we perform supervised fine-tuning (SFT) using the LLaMA-3 8B model. Our best-performing model strengthens the original LLaMA-3 8B Instruct model and achieves performance comparable to GPT-4o models on role-playing dialogue. We release our synthetic characters and instruction-tuning dialogues to support public research.

角色扮演大模型指令微调数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。