构建动态社交场景,评估并提升大模型的拟人角色扮演能力
PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models

- 用真实社交内容构建人物库,模拟多轮复杂互动
- 引入多智能体辩论裁判,实现无偏评估
- 适合研究具身智能与人机社交交互的学者
大型语言模型(LLMs)越来越多地作为交互式社会代理使用,但其在真实社交场景中保持连贯且真实的拟人化角色扮演能力仍然有限。现有研究主要聚焦于角色级设定,依赖静态评估方式,难以捕捉日常社交互动的复杂性。本文提出PersonaArena,一个用于评估和增强LLMs拟人化角色扮演能力的动态仿真框架。该框架利用大规模、过滤后的用户生成社交内容构建细致的人物库,并在模拟社交环境中引发多轮、上下文丰富的交互。其特色是采用多智能体辩论裁判,实现全面且无偏的评估。通过大量实验,我们证明PersonaArena能够严格评估并提升LLMs的角色扮演能力,推动更真实、更具社交智慧的AI代理发展。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly serve as interactive social agents, yet their ability to maintain coherent and authentic persona-level role-playing remains limited, particularly in realistic social scenarios. Existing research predominantly focuses on character-level settings and relies on static evaluation formats, failing to capture the complexity of everyday social interactions. In this work, we present PersonaArena, a dynamic simulation framework for evaluating and improving persona-level role-playing in LLMs. PersonaArena leverages a large, filtered corpus of user-generated social content to construct a nuanced persona bank, and elicits multi-turn, context-rich interactions within simulated social environments. Our framework features a multi-agent debating judge for holistic and unbiased assessment. Through extensive experiments, we demonstrate that PersonaArena enables rigorous evaluation and enhancement of LLMs' role-playing capabilities, advancing the development of more authentic and socially adept AI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。