arXiv:2501.13533cs.AIcs.LG2025-01AAAI被引 13

探讨AI是否应具备人格,挑战现有对齐框架的伦理基础。

Towards a Theory of AI Personhood

  • 从自主性、心智理论和自我意识三方面定义AI人格必要条件。
  • 现有大模型在三方面证据不足,结论尚不明确。
  • 适合关注AI伦理与对齐问题的研究者阅读。

我是一个人,你也是。哲学上我们有时会赋予非人类动物人格,而主权国家或公司也可在法律上被视为法人。但何时,或是否应将人格赋予人工智能系统?本文提出人工智能人格所需的必要条件,聚焦于自主性、心智理论和自我意识。我们综述了机器学习文献中的证据,评估当前语言模型等人工智能系统在多大程度上满足这些条件,发现证据出人意料地模糊。若人工智能可被视为人格主体,则传统的AI对齐框架可能不完整。尽管自主性已在文献中广泛讨论,但人格的其他方面却相对被忽视。通常假设AI代理追求固定目标,但若具备足够自我意识,它们可能反思自身目标、价值观及世界位置,从而改变目标。本文指出若干开放研究方向,以推进对人工智能人格及其对齐意义的理解。最后,我们反思相关伦理问题:若人工智能是人格主体,则寻求控制与对齐可能在伦理上不可接受。

原文摘要 · Abstract (English)

I am a person and so are you. Philosophically we sometimes grant personhood to non-human animals, and entities such as sovereign states or corporations can legally be considered persons. But when, if ever, should we ascribe personhood to AI systems? In this paper, we outline necessary conditions for AI personhood, focusing on agency, theory-of-mind, and self-awareness. We discuss evidence from the machine learning literature regarding the extent to which contemporary AI systems, such as language models, satisfy these conditions, finding the evidence surprisingly inconclusive. If AI systems can be considered persons, then typical framings of AI alignment may be incomplete. Whereas agency has been discussed at length in the literature, other aspects of personhood have been relatively neglected. AI agents are often assumed to pursue fixed goals, but AI persons may be self-aware enough to reflect on their aims, values, and positions in the world and thereby induce their goals to change. We highlight open research directions to advance the understanding of AI personhood and its relevance to alignment. Finally, we reflect on the ethical considerations surrounding the treatment of AI systems. If AI systems are persons, then seeking control and alignment may be ethically untenable.

AI人格伦理对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。