arXiv:2606.05553cs.CLcs.AI2026-06

评测角色扮演模型是否随剧情发展保持角色一致性,发现基于人物弧线的上下文效果最佳。

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?

论文配图:ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?
图 1 · 摘自论文原文
  • 构建基于人物心理轨迹的自动评测基准ArcANE,覆盖17部小说80个角色
  • 在未出现过的剧情场景中,使用人物弧线的上下文使模型表现提升显著
  • 微调后的ArcANE-8B/32B模型在未知情节中进一步放大优势,适合角色演进研究

角色扮演语言代理(RPLAs)应随故事推进动态展现角色价值观与行为的变化,而非维持固定人设。现有评测仅关注特定章节的事实记忆,无法衡量回应是否符合角色的心理演变轨迹,尤其在原文未涉及的情境中。本文提出自动构建的ArcANE(Arc-Aware Narrative Evaluation)评测基准,涵盖17部小说、80位主要角色。通过人物弧线将叙事划分为心理维度上的多个阶段,同一情境在不同阶段设置探针,覆盖原文内与外的情境。六种模型与六种上下文模式下,基于人物弧线的条件化策略在所有模型上均优于其他方法,且在原文未提及的情节中差距最大——此时检索无用。进一步在相同数据上微调开源模型,得到ArcANE-8B/32B,其在未知情境中的弧线优势更为明显。

原文摘要 · Abstract (English)

Role-playing language agents (RPLAs) should play characters whose values and behavior evolve as the story progresses, not maintain a fixed persona. Existing benchmarks measure factual recall at a given chapter, not whether responses align with the character's psychological trajectory, especially in scenarios the source text never explores. We introduce ArcANE (Arc-Aware Narrative Evaluation), an automatically constructed benchmark spanning 17 novels and 80 principal characters. A Character Arc segments the narrative into phases along a psychological axis, and each probe poses the same scenario across phases, spanning both situations within the source text and situations beyond it. Across six models and six context modes, conditioning on the Character Arc tops every other context strategy on every model, and the gap is largest on scenarios outside the source text where retrieval has nothing to find. We further fine-tune open-weight models on the same data to obtain ArcANE-8B/32B, which widen the Arc advantage even more on scenarios outside the source text.

角色扮演人物弧线评测基准语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。