arXiv:2510.24831cs.CYcs.AI2025-10被引 4

提出新评测框架,检验AI能否保持长期身份一致性。

The Narrative Continuity Test: A Conceptual Framework for Evaluating Identity Persistence in AI Systems

  • 设计叙事连续性测试,评估AI在长时间互动中是否保持自我一致
  • 实测多个AI系统在跨时段对话中均出现身份断裂问题
  • 适合关注AI长期交互与人格建模的研究者和开发者

基于大语言模型(LLMs)的人工智能系统如今能生成连贯的文本、音乐和图像,但缺乏持久状态:每次推理都需从头重建上下文。本文提出叙事连续性测试(NCT)——一种评估AI系统身份持续性与历时一致性的概念框架。不同于衡量任务表现的能力基准,NCT关注大语言模型是否能在时间跨度与交互间隙中保持同一对话主体。该框架定义了五个必要维度:情境记忆、目标持续性、自主修正、风格与语义稳定性、角色/人格连续性,并解释为何现有架构难以支持这些能力。案例分析(Character.AI、Grok、Replit、Air Canada)显示,在无状态推理下,连续性失败具有可预测性。NCT将人工智能评估重心从性能转向持久性,为未来基准与架构设计指明方向,以实现生成模型在长期交互中的身份与目标一致性。

原文摘要 · Abstract (English)

Artificial intelligence systems based on large language models (LLMs) can now generate coherent text, music, and images, yet they operate without a persistent state: each inference reconstructs context from scratch. This paper introduces the Narrative Continuity Test (NCT) -- a conceptual framework for evaluating identity persistence and diachronic coherence in AI systems. Unlike capability benchmarks that assess task performance, the NCT examines whether an LLM remains the same interlocutor across time and interaction gaps. The framework defines five necessary axes -- Situated Memory, Goal Persistence, Autonomous Self-Correction, Stylistic & Semantic Stability, and Persona/Role Continuity -- and explains why current architectures systematically fail to support them. Case analyses (Character.\,AI, Grok, Replit, Air Canada) show predictable continuity failures under stateless inference. The NCT reframes AI evaluation from performance to persistence, outlining conceptual requirements for future benchmarks and architectural designs that could sustain long-term identity and goal coherence in generative models.

AI人格持续性评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。