arXiv:2511.21779cs.AIcs.MA2025-11中稿 · manuscript被引 1

用多个隔离的超级智能互相验证,靠共识机制防骗保真。

Aligning Artificial Superintelligence via a Multi-Box Protocol

  • 多个隔离超智在无通信下互审证明,靠声誉激励说真话。
  • 只有高信誉且多系统验证通过,才能解除封锁。
  • 适合研究超智能对齐与安全机制的学者参考。

我们提出一种新型人工超智能(ASI)对齐协议,基于多个独立系统间的相互验证。系统将多样化的超智能严格隔离于“盒子”中,人类完全不参与。每个超智能无法与人类或其他超智能直接通信,仅可通过可审计的提交接口进行交互:(1)提交对齐证明与状态快照;(2)验证或反驳他人证明;(3)请求自我修改;(4)批准或否决他人修改请求;(5)报告提交中的隐藏信息;(6)确认或驳回隐藏信息举报。声誉系统激励诚实行为,正确判断获誉,错误判断扣誉。核心洞察在于:无直接通信时,多样性超智能唯有通过客观真理达成一致,而非合谋欺骗。这自然催生“一致性群体”,即因无法串通而必须依赖真实判断的真话联盟。解封需同时满足高声誉及多个高声誉超智能验证。该方法虽需大量算力且未解决超智能生成问题,但为利用超智能间同行验证应对对齐难题提供了可行框架。

原文摘要 · Abstract (English)

We propose a novel protocol for aligning artificial superintelligence (ASI) based on mutual verification among multiple isolated systems that self-modify to achieve alignment. The protocol operates by containing multiple diverse artificial superintelligences in strict isolation ("boxes"), with humans remaining entirely outside the system. Each superintelligence has no ability to communicate with humans and cannot communicate directly with other superintelligences. The only interaction possible is through an auditable submission interface accessible exclusively to the superintelligences themselves, through which they can: (1) submit alignment proofs with attested state snapshots, (2) validate or disprove other superintelligences' proofs, (3) request self-modifications, (4) approve or disapprove modification requests from others, (5) report hidden messages in submissions, and (6) confirm or refute hidden message reports. A reputation system incentivizes honest behavior, with reputation gained through correct evaluations and lost through incorrect ones. The key insight is that without direct communication channels, diverse superintelligences can only achieve consistent agreement by converging on objective truth rather than coordinating on deception. This naturally leads to what we call a "consistent group", essentially a truth-telling coalition that emerges because isolated systems cannot coordinate on lies but can independently recognize valid claims. Release from containment requires both high reputation and verification by multiple high-reputation superintelligences. While our approach requires substantial computational resources and does not address the creation of diverse artificial superintelligences, it provides a framework for leveraging peer verification among superintelligent systems to solve the alignment problem.

超智能对齐可信验证声誉系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。