arXiv:2502.00595cs.CLcs.AI2025-02被引 13

首个评估大模型玩文字角色扮演游戏的基准,测创意也测逻辑一致性。

RPGBENCH: Evaluating Large Language Models as Role-Playing Game Engines

  • 用结构化事件状态生成可玩游戏世界,确保逻辑连贯。
  • 多轮模拟交互中保持状态更新与规则约束,但长流程易出错。
  • 结合机器评分与大模型评判,兼顾客观规则和主观体验。

我们提出RPGBench,首个用于评估大语言模型(LLMs)作为文字角色扮演游戏(RPG)引擎的基准。该基准包含两大核心任务:游戏创建(GC)与游戏模拟(GS)。在GC中,模型需使用结构化事件-状态表示构建一个合法且可玩的游戏世界,保证逻辑一致性和正确终止条件;在GS中,模型需在多轮交互中持续更新状态并严格执行游戏规则。为全面评估性能,RPGBench融合客观与主观评价方法:客观指标验证事件机制遵循性及变量更新准确性,无需人工干预;主观指标如内容趣味性、动作质量与角色扮演能力,则通过“大模型作为裁判”框架,由强模型对输出进行评分。实证结果表明,当前顶尖模型虽能生成引人入胜的故事,但在复杂或长期场景中常难以维持一致且可验证的游戏机制。通过结合结构化规则评估与大模型评判,RPGBench为衡量大模型在文字RPG中平衡创造力、连贯性与复杂性的能力提供了新标准,推动更沉浸、可控的互动叙事发展。

原文摘要 · Abstract (English)

We present RPGBench, the first benchmark designed to evaluate large language models (LLMs) as text-based role-playing game (RPG) engines. RPGBench comprises two core tasks: Game Creation (GC) and Game Simulation (GS). In GC, an LLM must craft a valid and playable RPG world using a structured event-state representation, ensuring logical coherence and proper termination conditions. In GS, the LLM simulates interactive gameplay across multiple rounds while consistently updating states and enforcing game rules. To comprehensively assess performance, RPGBench integrates objective and subjective evaluation methodologies. Objective measures verify adherence to event mechanics and check variable updates without requiring human intervention. Subjective measures, such as content interestingness, action quality, and role-playing capability, are evaluated via an LLM-as-a-judge framework, where a strong LLM grades each candidate's outputs. Empirical results demonstrate that state-of-the-art LLMs can produce engaging stories but often struggle to implement consistent, verifiable game mechanics, particularly in long or complex scenarios. By combining structured, rule-based assessments with LLM-based judgments, RPGBench provides a new standard for evaluating how well LLMs can balance creativity, coherence, and complexity in text-based RPGs, opening avenues for more immersive and controllable interactive storytelling.

角色扮演大模型评测互动叙事

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。