arXiv:2504.18530cs.AIcs.CY2025-04NeurIPS被引 14

提出可扩展监督的量化框架,揭示弱模型监督强模型的成功概率规律。

Scaling Laws For Scalable Oversight

  • 将监督建模为能力不匹配者的博弈,用分段线性埃洛评分表征智能水平。
  • 在4种游戏中发现监督成功率随通用智能差距变化的规律,最大成功率达51.7%。
  • 理论推导多层嵌套监督最优层级数,适用于对强模型进行可信监督场景。

可扩展监督(scalable oversight)被提议作为控制未来超智能系统的关键策略,但其自身如何随规模变化仍不明确。为此,本文提出一个框架,将监督成功的概率建模为监督者与被监督系统能力的函数。该框架将监督视为能力不对称玩家间的博弈,使用特定于监督任务的埃洛分数,其值是通用智能的分段线性函数,包含任务无能和任务饱和两个平台。通过改进版的尼姆游戏验证框架后,将其应用于四种监督游戏:谋杀案(Mafia)、辩论(Debate)、后门代码(Backdoor Code)和兵棋推演(Wargames)。每种游戏均拟合出领域性能随通用人工智能能力变化的缩放规律。进一步开展嵌套可扩展监督(NSO)的理论研究,即可信模型监督不可信更强模型,后者成为下一层的可信模型。识别了NSO成功条件,并数值推导(部分解析)出最大化监督成功率的最优监督层级数。在通用埃洛差为400时,各游戏的NSO成功率分别为:谋杀案13.5%,辩论51.7%,后门代码10.0%,兵棋推演9.4%;监督更强系统时成功率进一步下降。

原文摘要 · Abstract (English)

Scalable oversight, the process by which weaker AI systems supervise stronger ones, has been proposed as a key strategy to control future superintelligent systems. However, it is still unclear how scalable oversight itself scales. To address this gap, we propose a framework that quantifies the probability of successful oversight as a function of the capabilities of the overseer and the system being overseen. Specifically, our framework models oversight as a game between capability-mismatched players; the players have oversight-specific Elo scores that are a piecewise-linear function of their general intelligence, with two plateaus corresponding to task incompetence and task saturation. We validate our framework with a modified version of the game Nim and then apply it to four oversight games: Mafia, Debate, Backdoor Code and Wargames. For each game, we find scaling laws that approximate how domain performance depends on general AI system capability. We then build on our findings in a theoretical study of Nested Scalable Oversight (NSO), a process in which trusted models oversee untrusted stronger models, which then become the trusted models in the next step. We identify conditions under which NSO succeeds and derive numerically (and in some cases analytically) the optimal number of oversight levels to maximize the probability of oversight success. We also apply our theory to our four oversight games, where we find that NSO success rates at a general Elo gap of 400 are 13.5% for Mafia, 51.7% for Debate, 10.0% for Backdoor Code, and 9.4% for Wargames; these rates decline further when overseeing stronger systems.

AI安全监督机制缩放定律嵌套监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。