构建可执行评测环境,评估大模型在互动叙事中的综合表现
NARRA-Gym for Evaluating Interactive Narrative Agents

- 设计NARRA-Gym环境,动态生成完整交互故事并记录全流程轨迹
- 九个前沿大模型表现差异显著,流畅性不等于用户体验与个性化能力
- 适合关注长程对话、用户自适应与角色扮演的AI研究者使用
互动叙事任务要求大语言模型在多轮交互中维持连贯且不断演进的故事,同时根据用户反馈进行调整。然而,现有评估方法多聚焦静态提示、孤立生成或事后评分,难以检验模型在故事生成、长上下文状态管理、节奏控制、角色模拟、共情个性化及故事相关产物生成等方面的综合能力。我们提出NARRA-Gym,一个可执行的评测环境,能将稀疏的情感种子转化为完整的交互式故事片段,并记录模型全周期行为轨迹,包括故事构建、记忆更新、规划、节奏干预及可选的产物生成。我们通过受控的LLM作为裁判评估九个前沿大模型,在八个基准人格设定下进行测试,并辅以人类评估,参与者对定制化输出进行评分。结果表明,模型在不同维度上表现差异显著:部分模型虽故事流畅,但在鲁棒性、用户体验或对抗敏感性的个性化方面仍存在缺陷。这说明互动叙事可作为评估长时序、用户自适应大模型行为的有效基准,超越单一故事质量。
原文摘要 · Abstract (English)
Interactive narrative tasks require LLMs to sustain a coherent, evolving story while adapting to a user over multiple turns. However, suitable benchmarks for this setting are limited: existing evaluations often focus on static prompts, isolated story generations, or post-hoc ratings, and therefore miss whether models can jointly manage story generation, long-context state and pacing, character simulation, empathic personalization, and story-grounded artifacts. We introduce NARRA-Gym, an executable evaluation environment that turns a sparse emotional seed into a complete interactive story episode and logs the full model-in-the-loop trajectory, including story construction, memory updates, planning, pacing interventions, and optional artifact synthesis. We evaluate nine frontier LLMs using a controlled LLM-as-judge sweep over eight benchmark personas and a human evaluation in which participants rate customized model outputs. Our results show substantial variation across models, personas, and evaluation dimensions: models that produce fluent stories can still fail on robustness, user experience, or resistance-sensitive personalization. These findings suggest that interactive narrative offers a useful benchmark for evaluating long-horizon, user-adaptive LLM behavior beyond isolated story quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。