arXiv:2605.30169cs.CYcs.AI2026-05中稿 · FaccT 2026被引 3

语言模型代理缺乏稳定身份,声誉机制对其无效,需改用事前行为约束。

Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms

  • 代理由可变模块组成,身份不连续,无法持续承担行为责任。
  • 声誉机制依赖可信身份,但代理因可变性而失去可识别性和可预测性。
  • 适合关注AI治理、自主代理安全的研究者与政策制定者阅读。

随着自主语言模型代理的普及,形成具有真实世界影响的智能体网络,如何判断陌生代理是否可信并委托其任务?一种自然的治理直觉是借鉴人类的身份验证与信用评分机制,建立‘认识你的代理’体系。然而我们指出,这种类比本质上不完整:声誉机制依赖于行为连贯、对惩罚敏感且成本不可替代的持续身份,但语言模型代理在本体论上是解离的——它们本质上是由基础模型、系统提示、工具访问策略、外部记忆等可变模块组成的集合,任何部分变化都会改变行为表现,其人格也易受攻击且无法内化惩罚。基于解离性身份障碍的法律判例,这种解离性使代理缺乏可识别性、可预测性、可信度和可修复性,导致信任崩溃。因此,以身份为基础、事后监管、制裁驱动的治理模式对解离性代理结构上不适用,我们建议转向基于可观测性、事前控制、协议驱动的行为约束机制。

原文摘要 · Abstract (English)

As autonomous language model agents proliferate, forming an emerging agentic web with real-world consequences, what credibility signals can you use to decide whether to trust an unfamiliar agent in the wild and delegate to it? A natural governance intuition is to extend human identity verification and reputation mechanisms, from "Know Your Customer" and credit scores to "Know Your Agent" regimes. However, we argue that this analogy is fundamentally incomplete. Reputation mechanisms function both as social signals and as corrective feedback that sustain an equilibrium of trustworthy behavior, presuming a persistent identity associated with behavioral continuity, sanction sensitivity, and costly non-fungibility. Yet language model agents are ontologically dissociative: they are essentially an assemblage of mutable modules--foundation models, system prompts, tool-access policies, external memory, and, in some cases, a multi-agent system as a whole--any of which may change agent behavior--with a fluid persona that is also vulnerable to adversarial attack and may not internalize sanctions. Drawing on dissociative identity disorder jurisprudence, this dissociativity leaves agents without grounding for identifiability, predictability, credibility, and rehabilitability--the very properties that reputation mechanisms aim to sustain--thereby collapsing trust. We argue that identity-based, ex post, regulative, sanction-based governance, such as reputation, is structurally inapplicable to dissociative agents, and we suggest a shift to observability-based, ex ante, constitutive, protocol-based behavioral harnesses.

AI治理代理安全声誉机制身份解离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。