提出动态会话评估框架,让角色扮演模型更持久地保持人设和对话质量。
DynSess: Dynamic Session-Level Evaluation and Optimization Framework for Role-Playing Agents

- 基于完整对话会话评分,捕捉长期行为表现。
- 新模型仅用少量参数就达到顶尖角色模型水平。
- 适合想提升角色一致性与长对话能力的研究者。
大型语言模型进行角色扮演本质上是会话级任务,要求代理在长时间多轮对话中维持角色身份与交互质量。然而现有评估与优化方法仍以单轮为单位,难以捕捉长程表现。本文提出 DynSess,一个统一的会话级框架。DynSess-Eval 通过针对长程行为的评分标准对完整对话会话进行打分;利用其会话级奖励信号,通过多轮前瞻搜索构建高质量训练轨迹,并训练出两个互补变体:DSPO(离策略)与 GSRPO(在线策略)。实验表明,DynSess-Eval 与人工评价高度一致,盲评结果显示,尽管参数量显著更少,DynSess-Character 在角色一致性与交互能力上仍媲美最强角色模型。数据集与代码将开源,推动后续研究。
原文摘要 · Abstract (English)
Role-playing with large language models is fundamentally a session-level task, requiring agents to sustain character identity and interaction quality across extended multi-turn conversations. Yet existing evaluation and optimization methods remain largely turn-level, failing to capture long-horizon quality. We propose DynSess, a unified session-level framework for role-playing agents. DynSess-Eval scores complete dialogue sessions via rubrics targeting long-horizon behaviors. Leveraging its session-level rewards, we construct high-quality training trajectories through multi-turn lookahead search and train DynSess-Character with two complementary variants: DSPO (off-policy) and GSRPO (on-policy). Experiments show that DynSess-Eval aligns with human judgments substantially better than prior evaluators, and blind human evaluation further shows that DynSess-Character matches the strongest character model despite using substantially fewer parameters, while maintaining strong role consistency and interactive ability. Our dataset and code will be released to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。