构建多任务角色扮演数据集,让AI更真实模仿人物语言风格
A Multi-Task Role-Playing Agent Capable of Imitating Character Linguistic Styles
- 设计新数据集MRstyle,包含7类任务和真实人物语录
- 提出StyleRPA模型,在7项任务上超越现有开源模型
- 适合需要角色化对话与文本生成的创作者和开发者
大型语言模型的兴起推动了角色扮演代理(RPAs)的发展,但现有RPAs主要关注角色基础属性的模拟,忽视语言风格的复现,且在多轮对话以外的任务中表现不佳,导致输出缺乏真实性。其根源在于现有角色数据集缺乏人物语录,仅限于多轮对话任务,限制了模型在其他任务中的表现。为此,我们构建了名为MRstyle的多任务角色扮演数据集,涵盖大量真实人物及其语录,并覆盖7种任务类型。基于此,我们提出了StyleRPA——一种多任务角色扮演代理(MRPA),在对话、字典、写作、故事生成、产品描述、音乐评论和开放问答共7项任务上显著优于近期开源大模型和基线方法。代码与数据将公开。
原文摘要 · Abstract (English)
The advent of large language models (LLMs) has significantly propelled the advancement of Role-Playing Agents (RPAs). However, current Role-Playing Agents predominantly focus on mimicking a character's fundamental attributes while neglecting the replication of linguistic style, and they are incapable of effectively replicating characters when performing tasks beyond multi-turn dialogues, which results in generated responses that lack authenticity. The reason current RPAs lack this capability is due to the nature of existing character datasets, which lack collections of character quotations and are limited to multi-turn dialogue tasks, constraining the RPA's performance across other task domains and failing to mimic a character's linguistic style. To address this gap, we developed a multi-task role-playing dataset named MRstyle, which encompasses a substantial number of real individuals along with their quotations and covers seven different tasks. On this basis, we develop StyleRPA, a Multi-Task Role-Playing Agent (MRPA) that significantly outperforms recent open-source LLMs and RPAs baselines on 7 tasks including Dialogue, Dictionary, Composition, Story Generation, Product Description, Music Commentary, and Open Question Answering. The code and data will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。