arXiv:2608.29798cs.CLcs.AI2026-08

让角色行为随任务动态调整,提升一致性与稳定性。

R$^2$A: Learning Persona Policies Through Persona Representation Learning and Runtime Alignment

论文配图:R$^2$A: Learning Persona Policies Through Persona Representation Learning and Runtime Alignment
图 1 · 摘自论文原文
  • 分两阶段学习角色策略:先建模角色特征,再实时对齐行为
  • 12个场景下性能优于基础模型和静态角色设定
  • 关键在角色表征学习,避免行为失衡,适合角色驱动任务

同一角色行为在不同情境中可能有益或有害,导致静态角色提取在多任务中表现不一。本文提出角色选择-实现框架,将行为生成建模为潜在角色状态,并分解为角色选择与角色实现两部分。静态角色提取与理想角色策略的差异分别定义为选择差距与实现差距。基于此框架,提出R$^2$A,一种两阶段角色策略学习方法。角色表征学习利用结构化‘谁—如何—做什么’描述编码目标角色的意图、条件行为原则及轨迹级表现;运行时对齐则移除显式角色指定,通过任务反馈联合校准行为选择与轨迹实现。在涵盖本文研究的四种可问责专业角色原则的12个评估设置中,R$^2$A整体优于基线模型与静态角色提取。消融实验表明,角色表征学习对防止运行时对齐产生行为失衡、实现更稳定的角色策略学习至关重要。

原文摘要 · Abstract (English)

The same Persona behavior can be beneficial in one context but harmful in another, causing static Persona elicitation to perform inconsistently across tasks. We introduce the Persona Selection--Realization Framework, which models behavior generation through a latent Persona state and decomposes it into Persona Selection and Persona Realization. The discrepancies between static Persona elicitation and an ideal Persona policy in these two components define the Selection Gap and Realization Gap, respectively. Building on this framework, we propose R$^2$A, a two-stage approach for learning Persona policies. Persona Representation Learning uses structured Who--How--What presentations to encode the target Persona's objective, conditional behavioral principles, and trajectory-level manifestations. Persona Runtime Alignment then removes the explicit Persona specification and jointly calibrates behavior selection and trajectory realization using task feedback. Across 12 evaluation settings covering the four principles of the Accountable-Professional Persona studied in this work, R$^2$A overall outperforms both the base model and static Persona elicitation. Ablation results further show that Persona Representation Learning is critical for preventing Runtime Alignment from producing behaviorally imbalanced policies and for achieving more stable Persona policy learning.

角色学习行为对齐策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。