arXiv:2412.16468cs.LG2024-12综述被引 17

探讨如何监管超越人类的超级智能,提出安全演进路径

The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment

  • 提出三种可扩展监督范式应对超智能监管难题
  • 分析现有方法在可能性与不可能性上的局限
  • 适合关注未来AI安全与可控发展的研究者

大型语言模型的出现引发了对人工超级智能(ASI)的讨论,这是一种假设性的人工智能系统,其智能水平将超越人类。尽管ASI仍属假设且远超当前人工智能能力,但探讨其潜在性、可行性及风险对未来发展至关重要。超级对齐概念源于可扩展监督,研究当直接人类监督不再可行时,如何监管日益强大的人工智能系统。本文聚焦超级对齐问题:‘对人工超级智能进行监督、控制与治理的过程’。首先回顾可扩展监督范式——夹心法、自我增强、弱到强泛化;随后从可能性与不可能性的视角分析现有范式的局限,探讨关键挑战,并提出未来人工智能系统安全持续改进的路径。

原文摘要 · Abstract (English)

The emergence of large language models (LLMs) has sparked discussion on Artificial Superintelligence (ASI), a hypothetical AI system that surpasses human intelligence. Although ASI remains hypothetical and far beyond current AI capabilities, discussing its potential and exploring its feasibility and potential risks is critical for the development of future AI systems. The idea of superalignment originates from scalable oversight, which studies how to supervise increasingly capable AI systems when direct human supervision becomes insufficient. In this paper, we focus on the superalignment problem: "The process of supervising, controlling, and governing artificial superintelligence." We first review scalable oversight paradigms-Sandwiching, Self-Enhancement, and Weak-to-Strong Generalization -- then analyze the limitations of current paradigms through the lens of possibility and impossibility, discuss key challenges, and propose pathways for the safe and continual improvement of future AI systems.

超级智能对齐问题可扩展监督安全演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。