梳理AI毁灭人类的可能路径,提出人类幸存的四种故事类型。
AI Survival Stories: a Taxonomic Analysis of AI Existential Risk
- 构建基于两大前提的生存故事分类框架。
- 指出不同路径下人类幸存的概率差异,估算毁灭概率。
- 为应对AI风险提供针对性策略建议,适合政策制定者参考。
自ChatGPT发布以来,关于人工智能是否对人类构成生存威胁的讨论持续升温。本文提出一个通用框架,用于分析人工智能的生存性风险。核心围绕两个前提展开:其一,人工智能系统将变得极其强大;其二,若人工智能系统极度强大,将毁灭人类。基于这两个前提,我们构建了人类在遥远未来得以幸存的生存故事分类体系。在每一种故事中,至少一个前提不成立:或科学障碍阻止AI达到极致强大;或人类主动禁止相关研究;或超级智能因目标设定而不会毁灭人类;或我们能可靠检测并关闭具有毁灭意图的系统。本文指出不同生存故事面临不同挑战,并引导出相应应对策略。最后,利用该分类体系,对人类被AI毁灭的概率(P(doom))给出初步估算。
原文摘要 · Abstract (English)
Since the release of ChatGPT, there has been a lot of debate about whether AI systems pose an existential risk to humanity. This paper develops a general framework for thinking about the existential risk of AI systems. We analyze a two premise argument that AI systems pose a threat to humanity. Premise one: AI systems will become extremely powerful. Premise two: if AI systems become extremely powerful, they will destroy humanity. We use these two premises to construct a taxonomy of survival stories, in which humanity survives into the far future. In each survival story, one of the two premises fails. Either scientific barriers prevent AI systems from becoming extremely powerful; or humanity bans research into AI systems, thereby preventing them from becoming extremely powerful; or extremely powerful AI systems do not destroy humanity, because their goals prevent them from doing so; or extremely powerful AI systems do not destroy humanity, because we can reliably detect and disable systems that have the goal of doing so. We argue that different survival stories face different challenges. We also argue that different survival stories motivate different responses to the threats from AI. Finally, we use our taxonomy to produce rough estimates of P(doom), the probability that humanity will be destroyed by AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。