arXiv:2604.08986cs.CLcs.AI2026-04

让大模型在保持任务能力的同时,更忠实地扮演不同人设。

PerMix-RLVR: Preserving Persona Expressivity under Verifiable-Reward Alignment

论文配图:PerMix-RLVR: Preserving Persona Expressivity under Verifiable-Reward Alignment
图 1 · 摘自论文原文
  • 训练时混合人设并用可验证奖励强化学习,减少对提示的敏感性。
  • 在MATH500上人设稳定性提升21.2%,在PersonaGym上人设忠实度提升11.4%。
  • 适合需要稳定表现又需精准角色扮演的应用场景。

人格化提示被广泛用于引导大语言模型行为并提升其指令响应能力,但最优人格设定的识别耗时且对输出质量的影响尚不明确。以往工作主要在推理阶段通过提示策略优化,增加计算开销。本文提出在训练阶段解决人格敏感性问题,使模型能适应多样人格同时保持任务性能。我们发现基于可验证奖励的强化学习(RLVR)虽提升了对有害人格变化的鲁棒性,却会削弱角色扮演中的表达忠实度。为此提出PerMix-RLVR,一种人格混合的强化学习策略,在保留强鲁棒性的同时增强人格忠实度。实验显示,该方法在MATH500上将人设稳定性分数(PSS)提升21.2%,在PersonaGym上人设忠实度提升11.4%。

原文摘要 · Abstract (English)

Persona prompting has been widely adopted to steer large language models (LLMs) behavior and improve their instruction performance by assigning specific characters. However, identifying an optimal persona is time-consuming, and its impact on output quality remains poorly understood. Prior work has mainly addressed this issue at the prompt level via inference-time strategies, incurring additional computation. In this work, we avoid inference-time prompt search by tackling persona sensitivity during training, aiming to train models that adapt their behavior to diverse personas while preserving task performance. In particular, we find that reinforcement learning with verifiable rewards (RLVR) systematically reduces sensitivity to persona prompts, but also reveals an inherent trade-off of outcome-based optimization: while RLVR improves robustness on tasks with verifiable goals, it can also degrade persona expressivity when needed, e.g., in-character role-playing. To address this limitation, we propose PerMix-RLVR, a persona-mixed RLVR strategy that mitigates the persona robustness-fidelity trade-off, preserving strong robustness to harmful persona variation while enabling faithful persona adoption when required. Concretely, PerMix-RLVR improves persona stability score (PSS) over RLVR by +21.2% on MATH500, while also enhancing persona fidelity by +11.4% on PersonaGym.

大模型人格建模强化学习角色扮演

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。