构建可动态演化的角色状态树,让长对话角色更真实连贯。
PHASE-Tree: Modeling Character-State Evolution in Long-Horizon Role-Playing Dialogue

- 用多时序树结构分层管理角色身份与状态,支持局部更新不破坏整体一致性。
- 在长对话任务中提升角色表现19.7%~15.1%,显著优于现有方法。
- 适合需要长期角色演化的游戏、叙事生成等场景研究者使用。
长时段角色扮演要求角色在剧情推进中保持辨识度。现有工作存在两大缺陷:角色表示通常为静态画像,难以局部更新而不影响不变特质;评估基准多聚焦于人格保留和记忆召回,而非模型是否基于角色当前演化状态说话。本文提出PHASE-Tree,一种具有不可变身份根节点的多时序角色状态树,包含可变的人格、会话和瞬时层,使每个可变字段可独立定位并实现跨剧集或剧中更新。生成通过显式文本输入或隐式参数适配实现。为衡量演化态生成,引入LongEvoRoleBench,结合四组长对话语料(跨剧集演化)与四组短对话语料(场景内状态追踪),统一采用下一语句预测协议。在长对话核心任务上,文本式PHASE-Tree在12项指标中11项排名第一,相较内部变体提升角色级、语义与嵌入分数19.7%、12.4%、15.1%。双盲200次响应测试中,人工评分与GPT-4.1判别相关系数达0.65;在描述性提示子集上,总体差异为+0.20。该优势在不同大模型裁判及生成基线中持续存在。
原文摘要 · Abstract (English)
Long-horizon role-playing demands that characters remain recognizable as they evolve with the narrative. Yet existing work falls short on two fronts: representations are typically static profiles that cannot be updated locally without destabilizing unchanged traits, and benchmarks mainly test persona preservation and memory recall rather than whether a model speaks from a character's currently evolved state. We address both. PHASE-Tree is a multi-timescale character-state tree with an immutable identity root and mutable persona, session, and moment layers, making each mutable field an addressable target for localized within- and cross-episode updates. It conditions generation through explicit textual provision or implicit parametric adaptation. To measure evolved-state generation, we introduce LongEvoRoleBench, which pairs four long-dialogue corpora for cross-episode evolution with four short-dialogue corpora as within-scene state-tracking checks, under a unified next-utterance protocol. On the long-dialogue core, textual PHASE-Tree ranks first in 11 of 12 dataset-metric cells against internal variants and all 12 cells against external textual baselines, improving character-level, semantic, and embedding scores by 19.7%, 12.4%, and 15.1% respectively. In a blinded 200-response study, human ratings correlate with the GPT-4.1 judge (Pearson r= 0.65); on descriptive n= 10 PT and NR prompt subsets, the Overall difference is +0.20. The long-dialogue Sem advantage persists across LLM judges and generation backbones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。