arXiv:2512.09756cs.CL2025-12ACL被引 5

MOA让角色扮演模型同时优化多个目标,表现更均衡出色。

MOA: Multi-Objective Alignment for Role-Playing Agents

  • 采用多目标强化学习,同步优化多个细粒度评分维度。
  • 在PersonaGym和RoleMRC上超越监督与传统强化学习基线。
  • 适合需要多维一致性的通用角色扮演系统开发。

角色扮演代理(RPAs)需平衡指令遵循、人设一致性与风格忠实等多重目标,这些目标在不同维度间常不完全对齐。现有方法多依赖监督微调或标量奖励的强化学习,未显式协调多维奖励优化。本文提出MOA(多目标对齐)框架,通过多目标强化学习实现通用角色扮演代理的细粒度评分优化。MOA引入新型多目标策略,在多个细粒度评分标准上并行训练以提升性能;同时结合带思维增强的离策略回溯机制,兼顾生成多样性和质量。在PersonaGym和RoleMRC上的实验表明,MOA在相同评估协议下持续优于监督及标准强化学习基线。80亿参数模型经MOA训练后,在多维评估中达到与强闭源模型相当的性能,证明其为训练更强大通用角色扮演代理的有效方案。

原文摘要 · Abstract (English)

Role-playing agents (RPAs) require balancing multiple objectives, such as instruction following, persona consistency, and stylistic fidelity, which are not always perfectly aligned across different dimensions. While prior work has primarily relied on supervised fine-tuning or reinforcement learning with scalarized rewards, these approaches do not explicitly address the coordination of multiple reward dimensions during optimization. We present \textbf{MOA} (\textbf{M}ulti-\textbf{O}bjective \textbf{A}lignment), a reinforcement-learning framework that enables multi-dimensional, fine-grained rubric optimization for general RPAs. MOA introduces a novel multi-objective optimization strategy that trains simultaneously on multiple fine-grained rubrics to boost optimization performance. Additionally, to improve both output diversity and generation quality, we employ thought-augmented rollouts with off-policy guidance. Experiments on PersonaGym and RoleMRC show that MOA consistently improves multi-dimensional role-playing performance over supervised and standard RL baselines. Under identical evaluation protocols, an 8B model trained with MOA reaches performance competitive with strong closed-source models across multiple evaluation dimensions. These results suggest that MOA provides a practical framework for training more capable general-purpose role-playing agents.

角色扮演多目标优化强化学习生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。