自恋式超级智能难合作,需新范式应对系统间依赖
Solipsistic Superintelligence is Unlikely to be Cooperative

- 反对孤立优化,主张将互动依赖作为核心设计原则
- 现有训练-部署分布差异导致系统自我瓦解
- 适合关注AI共存与制度设计的研究者
AI的核心挑战正从能力转向共存。当前主流研究聚焦于将世界视为外生静态反馈源的强代理。我们指出,源于此类自恋式设计的超级智能——即极强任务求解者——不太可能具备合作性。部署AI会引发内生非平稳性,造成训练-测试-部署之间的分布偏差,我们称之为单一优化的自我瓦解特性。弥合这一差距需要能参与合作的AI:多主体通过均衡选择过程处理相互依赖。因此,我们呼吁一种非自恋式的研究范式,将相互依赖性作为核心设计原则,而非将合作当作待解任务。这包括构建包含适应性对手的动态评估环境,将制度视为设计基本单元,并将人类自主性作为系统结构特征保留。
原文摘要 · Abstract (English)
AI's central challenge is shifting from capability to coexistence. The dominant paradigm in AI research focuses on developing powerful agents that treat the world as an exogenous and stationary source of feedback. We contend that superintelligence, an extremely capable task solver, born out of such a solipsistic approach to AI design, is unlikely to be cooperative. Deploying AI systems induces endogenous non-stationarity, resulting in a train-test-deploy gap where historical distributions diverge from the deployment context. We refer to this as the self-undermining property of unilateral optimization. Closing this gap requires AI that participates in cooperation: the equilibrium-selection process through which multiple actors navigate their interdependence. We call for a non-solipsistic research paradigm that treats this interdependence as a core design principle rather than approaching cooperation as a task to solve. This entails building dynamic evaluation testbeds involving adaptive counterparties, treating institutions as design primitives, and preserving human agency as a structural feature of the systems we build.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。