arXiv:2504.03731cs.AI2025-04中稿 · ICLR被引 1

建立可扩展监督的评估基准,量化人类反馈机制对说真话的优势。

A Benchmark for Scalable Oversight Protocols

  • 基于代理得分差(ASD)设计评估框架,衡量反馈机制对真相的激励程度。
  • 首次提供可复现的基准测试工具包,支持快速对比不同监督协议性能。
  • 适用于研究对齐、可信AI的学者,尤其关注可扩展人类反馈机制者。

随着人工智能代理超越人类能力,可扩展监督——即如何有效向可能超人类的AI模型提供人类反馈——成为确保对齐的关键挑战。尽管已有多种可扩展监督协议被提出,但缺乏系统性的实证评估框架进行比较。现有研究虽尝试实验性分析如「辩论」等协议,但其实验设计难以推广至其他协议。本文提出可扩展监督基准,基于代理得分差(ASD)这一指标,衡量反馈机制在促进说真话与压制欺骗方面的有效性。我们提供一个Python工具包,支持快速且可竞争地在该基准上评估各类协议,并以辩论协议为例开展演示实验。

原文摘要 · Abstract (English)

As AI agents surpass human capabilities, scalable oversight -- the problem of effectively supplying human feedback to potentially superhuman AI models -- becomes increasingly critical to ensure alignment. While numerous scalable oversight protocols have been proposed, they lack a systematic empirical framework to evaluate and compare them. While recent works have tried to empirically study scalable oversight protocols -- particularly Debate -- we argue that the experiments they conduct are not generalizable to other protocols. We introduce the scalable oversight benchmark, a principled framework for evaluating human feedback mechanisms based on our agent score difference (ASD) metric, a measure of how effectively a mechanism advantages truth-telling over deception. We supply a Python package to facilitate rapid and competitive evaluation of scalable oversight protocols on our benchmark, and conduct a demonstrative experiment benchmarking Debate.

可扩展监督对齐研究评估基准人类反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。