arXiv:2502.04675cs.AIcs.CL2025-02被引 5

用递归自我批评让超人智能可被可靠监督

Scalable Oversight for Superhuman AI via Recursive Self-Critiquing

  • 让AI不断批评论文,层级越高越容易判断对错
  • 实验显示高阶批判比直接评判更准确可靠
  • 适合需要大规模监督的超智能系统研发

随着AI在复杂任务中的能力持续超越人类,现有对齐技术(如SFT和RLHF)面临根本性挑战:依赖人类直接评估的方法在AI输出超出人类认知阈值时变得不可行。为此,本文提出两个假设:(1) 对批评的批评比批评本身更容易——将验证易于生成的普遍规律扩展到批评领域,因批评本身就是一种特殊生成;(2) 此难度关系具有递归性,当直接评估不可行时,进行更高阶的批判(如三重批评)能提供更可行的监督路径。我们通过人类-人类、人类-AI和AI-AI三类实验,验证了递归自我批评在可扩展AI监督中的潜力。结果表明,递归批判是一种极具前景的规模化监督方法。

原文摘要 · Abstract (English)

As AI capabilities increasingly surpass human proficiency in complex tasks, current alignment techniques, including SFT and RLHF, face fundamental challenges in ensuring reliable oversight. These methods rely on direct human assessment and become impractical when AI outputs exceed human cognitive thresholds. In response to this challenge, we explore two hypotheses: (1) \textit{Critique of critique can be easier than critique itself}, extending the widely-accepted observation that verification is easier than generation to the critique domain, as critique itself is a specialized form of generation; (2) \textit{This difficulty relationship holds recursively}, suggesting that when direct evaluation is infeasible, performing higher-order critiques (e.g., critique of critique of critique) offers a more tractable supervision pathway. We conduct Human-Human, Human-AI, and AI-AI experiments to investigate the potential of recursive self-critiquing for AI supervision. Our results highlight recursive critique as a promising approach for scalable AI oversight.

AI对齐递归批判监督机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。