arXiv:2502.11387cs.CL2025-02ACL被引 22

构建细粒度角色扮演与指令遵循评测基准,提升大模型角色一致性与指令执行能力。

RoleMRC: A Fine-Grained Composite Benchmark for Role-Playing and Instruction-Following

  • 设计三类角色扮演任务:多轮对话、机器阅读理解、嵌套复杂指令。
  • 包含10.2k角色元数据、37.9k合成指令、1.4k测试样本,支持精细评估。
  • 适合作为大模型角色扮演与指令遵循能力训练与评测的基准工具。

角色扮演对大语言模型(LLMs)在保持角色身份与能力边界的同时遵循多样化指令至关重要。现有数据集多关注角色风格与知识范围控制,却忽视了角色扮演中的指令遵循场景。本文提出细粒度角色扮演与指令遵循复合基准RoleMRC,包含:(1) 理想角色与人类间的多轮对话,涵盖自由聊天或基于文本段落的讨论;(2) 角色扮演下的机器阅读理解,根据段落可答性与角色能力进行回答、拒绝或尝试;(3) 嵌套、多轮且有优先级的复杂指令场景。最终的RoleMRC包含10.2k角色元数据池、37.9k高质量合成角色扮演指令及1.4k测试样本。我们构建了评估流水线,量化评估主流LLMs及其在该数据上微调后的表现。跨数据集验证表明,基于RoleMRC微调的模型在提升指令遵循能力的同时,未损害通用角色扮演与推理能力。我们还分析了微调后模型在神经层面的激活模式。

原文摘要 · Abstract (English)

Role-playing is important for Large Language Models (LLMs) to follow diverse instructions while maintaining role identity and the role's pre-defined ability limits. Existing role-playing datasets mostly contribute to controlling role style and knowledge boundaries, but overlook role-playing in instruction-following scenarios. We introduce a fine-grained role-playing and instruction-following composite benchmark, named RoleMRC, including: (1) Multi-turn dialogues between ideal roles and humans, including free chats or discussions upon given passages; (2) Role-playing machine reading comprehension, involving response, refusal, and attempts according to passage answerability and role ability; (3) More complex scenarios with nested, multi-turn and prioritized instructions. The final RoleMRC features a 10.2k role profile meta-pool, 37.9k well-synthesized role-playing instructions, and 1.4k testing samples. We develop a pipeline to quantitatively evaluate the fine-grained role-playing and instruction-following capabilities of several mainstream LLMs, as well as models that are fine-tuned on our data. Moreover, cross-evaluation on external role-playing datasets confirms that models fine-tuned on RoleMRC enhances instruction-following without compromising general role-playing and reasoning capabilities. We also probe the neural-level activation maps of different capabilities over post-tuned LLMs. Access to our RoleMRC, RoleMRC-mix and Codes: https://github.com/LuJunru/RoleMRC.

角色扮演指令遵循评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。