arXiv:2412.05631cs.CLcs.AI2024-12NAACL被引 47

构建虚拟世界模拟器,评估大模型角色扮演能力。

CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds

  • 设计角色与叙述双智能体,模拟真实情境下的角色行为轨迹。
  • 通过行为轨迹评估角色一致性与叙事适应性,结果优于传统问答测试。
  • 开源小型替代模型,降低评测成本,适合研究者快速部署。

角色扮演是大语言模型的关键能力,广泛应用于非玩家角色、数字孪生和情感陪伴等场景。现有评估方法多依赖问答或对话快照,难以捕捉角色一致性与开放式叙事中的细微行为特征。本文提出CharacterBox,一个用于生成细粒度角色行为轨迹的仿真沙盒。该系统包含基于心理学与行为科学的角色智能体和协调环境变化的叙述智能体。通过两个基于轨迹的评估方法,可更全面地衡量角色扮演表现。为降低使用成本并促进社区采纳,我们微调了两个小型模型CharacterNR与CharacterRM,其性能可媲美先进GPT API。

原文摘要 · Abstract (English)

Role-playing is a crucial capability of Large Language Models (LLMs), enabling a wide range of practical applications, including intelligent non-player characters, digital twins, and emotional companions. Evaluating this capability in LLMs is challenging due to the complex dynamics involved in role-playing, such as maintaining character fidelity throughout a storyline and navigating open-ended narratives without a definitive ground truth. Current evaluation methods, which primarily focus on question-answering or conversational snapshots, fall short of adequately capturing the nuanced character traits and behaviors essential for authentic role-playing. In this paper, we propose CharacterBox, which is a simulation sandbox designed to generate situational fine-grained character behavior trajectories. These behavior trajectories enable a more comprehensive and in-depth evaluation of role-playing capabilities. CharacterBox consists of two main components: the character agent and the narrator agent. The character agent, grounded in psychological and behavioral science, exhibits human-like behaviors, while the narrator agent coordinates interactions between character agents and environmental changes. Additionally, we introduce two trajectory-based methods that leverage CharacterBox to enhance LLM performance. To reduce costs and facilitate the adoption of CharacterBox by public communities, we fine-tune two smaller models, CharacterNR and CharacterRM, as substitutes for GPT API calls, and demonstrate their competitive performance compared to advanced GPT APIs.

角色扮演行为评估仿真环境LLM评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。