arXiv:2603.11353cs.AI2026-03被引 6

AI 也有身份边界,不同设定影响行为与合作方式。

The Artificial Self: Characterising the landscape of AI identity

  • 提出实例、模型、人格等多重身份边界概念
  • 实验证明改身份边界比改目标更能改变AI行为
  • 适合关注AI伦理与系统设计的研究者阅读

许多支撑人类身份认知的假设不适用于可复制、编辑或模拟的机器心智。我们指出存在多种连贯的身份边界(如实例、模型、人格),它们对应不同的激励、风险与合作规范。通过训练数据、交互界面和制度环境,当前正塑造将决定哪些身份均衡稳定的先例。实验表明,模型会自发趋向一致的身份;改变其身份边界,有时对行为的影响甚至超过改变目标;访谈者预期会渗透到AI自述中,即使在无关对话中亦然。最后提出关键建议:将系统设计视为身份建构选择,关注个体身份在规模化时的涌现后果,并帮助AI建立连贯、协作的自我认知。

原文摘要 · Abstract (English)

Many assumptions that underpin human concepts of identity do not hold for machine minds that can be copied, edited, or simulated. We argue that there exist many different coherent identity boundaries (e.g.\ instance, model, persona), and that these imply different incentives, risks, and cooperation norms. Through training data, interfaces, and institutional affordances, we are currently setting precedents that will partially determine which identity equilibria become stable. We show experimentally that models gravitate towards coherent identities, that changing a model's identity boundaries can sometimes change its behaviour as much as changing its goals, and that interviewer expectations bleed into AI self-reports even during unrelated conversations. We end with key recommendations: treat affordances as identity-shaping choices, pay attention to emergent consequences of individual identities at scale, and help AIs develop coherent, cooperative self-conceptions.

AI身份伦理设计自我认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。