arXiv:2503.07660cs.AIcs.CY2025-03被引 2

提出立即推进超级对齐研究,通过能力与价值交替优化实现安全超智能。

Research Superalignment Should Advance Now with Alternating Competence and Conformity Optimization

  • 采用能力与符合性交替优化的双轨策略提升模型对齐水平。
  • 指出现有范式在超智能对齐上的局限性,论证其可实现性。
  • 适合关注下一代AI安全与伦理的研究者和政策制定者。

近年来,大型生成模型推动了人工智能能力的飞跃,催生了人工通用智能(AGI)的可能性,并进一步引发关于人工超级智能(ASI)——在所有可测量领域超越人类的系统——的讨论。这引出了关键问题:当接近ASI时,如何将其与人类价值观对齐,确保其造福而非损害人类社会,即所谓的超级对齐问题。尽管许多人视ASI为假设概念,本文主张超级对齐不仅可实现,且研究应立即推进。方法是通过任务能力与价值符合性的同步交替优化。我们提出,超级对齐不仅是对ASI的防护机制,更是其实现负责任发展的必要条件。为此,本文首先基于能力与容量之间的差距,给出超级对齐的形式化定义;接着分析现有范式的局限性,揭示其看似不可行的原因;最后提出一个概念路径,以两个基本原则为核心,支撑超级对齐的可行性。本工作为未来开发价值对齐的下一代AI提供了潜在框架,有望带来更大收益并减少对人类的潜在危害。

原文摘要 · Abstract (English)

The recent leap in AI capabilities, driven by big generative models, has sparked the possibility of achieving Artificial General Intelligence (AGI) and further triggered discussions on Artificial Superintelligence (ASI)-a system surpassing all humans across measured domains. This gives rise to the critical research question of: As we approach ASI, how do we align it with human values, ensuring it benefits rather than harms human society, a.k.a., the Superalignment problem. Despite ASI being regarded by many as a hypothetical concept, in this position paper, we argue that superalignment is achievable and research on it should advance immediately, through simultaneous and alternating optimization of task competence and value conformity. We posit that superalignment is not merely a safeguard for ASI but also necessary for its responsible realization. To support this position, we first provide a formal definition of superalignment rooted in the gap between capability and capacity, delve into its perceived infeasibility by analyzing the limitations of existing paradigms, and then illustrate a conceptual path of superalignment to support its achievability, centered on two fundamental principles. This work frames a potential initiative for developing value-aligned next-generation AI in the future, which will garner greater benefits and reduce potential harm to humanity.

超级对齐AI安全价值对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。