arXiv:2607.27816cs.CLcs.AI2026-07

用真人模拟器评估角色扮演模型,更真实地反映用户个性化体验。

Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation

论文配图:Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
图 1 · 摘自论文原文
  • 构建用户模拟器,让角色扮演模型与虚拟用户多轮互动
  • 基于300个角色档案,生成个性化评价标准,提升评测准确性
  • 可分析每个用户与模型的互动细节,适合精细化评估

角色扮演代理(RPAs)已成为大语言模型的重要消费应用。用户通过多轮对话与RPAs互动,以获得情感慰藉等体验,因此可靠评估对衡量能力、系统比较和改进至关重要。现有基准通常要求模型延续固定对话历史,并使用脱离用户的固定评分标准进行评估。我们实证发现两种缺陷:第一,模型输出受历史对话影响,难以科学评估其在真实多轮场景中的角色扮演能力;第二,用户体验因人而异,传统固定评分标准未必反映用户满意度。为此,我们提出PALATE(Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation),一个基于用户模拟器的可扩展基准。PALATE包含300个角色档案,主评估阶段为每位用户训练5个模拟器,使其与候选模型在预冻结的角色档案上进行自由形式的多轮对话。除通用质量评分外,还构建个性化评分标准,实验显示其与人工判断的一致性高于通用标准。在16个候选模型的主评估中,PALATE分别刻画了通用回合质量、长程会话能力以及每位用户的多轮交互体验,提供可解释的用户-模型配对评估,而非将系统压缩为单一用户无关排名。

原文摘要 · Abstract (English)

Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA's output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles. Alongside a general quality rubric, we construct personalized rubrics to measure user satisfaction; on held-out annotated data, the personalized rubrics show higher agreement with human judgments than the general rubric. In the main evaluation of 16 candidates, PALATE separately characterizes generic turn quality, long-horizon session capability, and per-user experience on multi-turn trajectories co-constructed by each candidate. It thereby produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.

角色扮演用户模拟评估基准多轮对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。