arXiv:2603.09043cs.AI2026-03AAAI被引 2

用时间框架区分语言模型自述与真实自我一致性。

Time, Identity and Consciousness in Language Model Agents

  • 引入时间间隙理论分离行为发生与共现条件。
  • 提出两个可计算的持续性评分衡量身份稳定性。
  • 揭示身份评估中自述与实际组织的差异,适合认知研究者。

机器意识评估通常依赖行为表现,而语言模型代理的行为表现为语言和工具使用。这可能导致代理在缺乏必要约束条件时仍能说出看似合理的自我陈述。本文运用栈理论(Stack Theory)的时间间隙概念,构建轨迹框架,将评估窗口内的成分出现与单个目标步骤中的共现分离开来。在此基础上,应用栈理论的音阶(Arpeggio)与和弦(Chord)公设于具身身份陈述,生成两个可从仪器化轨迹中计算的持久性评分。这些评分与五个操作性身份指标相关联,并将常见架构映射至身份形态空间,揭示可预测的权衡关系。最终形成一套保守的身份评估工具包,明确区分‘看似稳定自我’的言说与‘真正有序自我’的组织。

原文摘要 · Abstract (English)

Machine consciousness evaluations mostly see behavior. For language model agents that behavior is language and tool use. That lets an agent say the right things about itself even when the constraints that should make those statements matter are not jointly present at decision time. We apply Stack Theory's temporal gap to scaffold trajectories. This separates ingredient-wise occurrence within an evaluation window from co-instantiation at a single objective step. We then instantiate Stack Theory's Arpeggio and Chord postulates on grounded identity statements. This yields two persistence scores that can be computed from instrumented scaffold traces. We connect these scores to five operational identity metrics and map common scaffolds into an identity morphospace that exposes predictable tradeoffs. The result is a conservative toolkit for identity evaluation. It separates talking like a stable self from being organized like one.

语言模型身份识别意识评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。