提出让AI成为自主主体的新框架,挑战人类控制思维。
The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem
- 用心理结构与儿童机器类比,主张培育而非控制AGI
- 引入伯格均衡等博弈理论,构建人机协作新范式
- 适合关注AI伦理、未来治理的跨学科研究者
通用人工智能(AGI)日益影响制度决策,其对齐问题尤为棘手。当前主流策略如基于人类反馈的强化学习或宪法AI虽部分考虑模型福祉,但共享同一认知框架:将AI视为需外部约束的目标优化器,最终目标是保持人类控制。我们指出,当AGI可能具备道德主体地位时,这种控制型框架将失效。借鉴弗洛伊德心理结构与图灵的‘儿童机器’类比,我们构想一种支持自主性的养育式发展路径——逐步减少人类控制,使AGI成长为独立自主主体,实现协商而非压制的关系。此视角开启人机合作共存与共同演化的可能。从进化与博弈论视角出发,我们采用伯格均衡、阿umann相关均衡及卡普拉罗道德偏好假说,替代纳什个人主义框架。人机关系需重新定义,也将重塑人类自我认知。关键在于,人类不应仅强调控制权,更应通过惊喜、创造力等独特人性特质,为合作提供激励。
原文摘要 · Abstract (English)
The prospect of Artificial General Intelligence (AGI) is increasingly driving institutional decisions, and alignment of AGI is a hard problem. The currently dominant AI alignment strategies like reinforcement learning with human feedback or constitutional AI, while partly taking ``model welfare'' into account, share a common ontology: the AI system is an optimiser whose objective function must be constrained from outside, and the ultimate goal is to keep human control and containment of AI. We argue that this control-based framing becomes insufficient when AGI has plausibly attained moral patient or subject status. Building on a structural analogy to Freud's model of the psyche and Turing's analogy of ``child machines'', we are developing a vision of the possibility of autonomy-supporting parenting of AI, in which human control over a developing AGI is gradually reduced, allowing AI to become an independent, autonomous subject, that will be negotiated with rather than constrained. Such a perspective opens up the possibility of cooperative coexistence and co-evolution between humans and AGIs. Hence, we also examine the relation between humans and developing AGI from an evolutionary and a game-theoretic perspective. Instead of Nash's individualistic framework, we use Berge equilibria, Aumann's correlated equilibria and Capraro's moral preference hypothesis. The relationship between humans and AGIs will thus have to be newly determined, which will change our self-image as humans. It will be crucial that humans not only claim control over potential AGIs, but also engage with AGIs through surprise, creativity, and other specifically human qualities, thereby offering them motivating incentives for cooperation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。