arXiv:2412.16238cs.AIcs.LG2024-12

用代数方法评估裁判表现,比多数投票更准更可靠。

Algebraic Evaluation Theorems

  • 基于错误独立性假设,通过批量决策数据计算裁判准确率
  • 三名以上裁判即可精确估计其正确率,且可给出置信区间
  • 适合用于无法理解任务的AI系统评估与自动终止监控链

多数投票(MV)是群体智慧的经典算法。1785年康多塞的陪审团判决定理揭示了MV在特定条件下最优。同样的错误独立性假设可用于证明陪审团评估定理(AE),实现对裁判表现的纯代数评估。三个或更多二元裁判即可确定其在测试中的唯一两种正确性统计量。相比MV,AE有三大优势:第一,假设更宽松,能处理准确率低于50%的裁判;第二,基于误差独立性具有点状精度,支持多准确率方法,提升标注准确率并附带经验不确定性边界;第三,能自检误差独立性假设是否失效。使用美国社区调查的人口数据实验验证了AE在实践中的优越性。论文还讨论了两个对AI安全的启示:解决无限监控链的终止问题(谁来评判评判者?)以及超级对齐难题(如何评估执行我们不理解任务的智能体?)。

原文摘要 · Abstract (English)

Majority voting (MV) is the prototypical ``wisdom of the crowd'' algorithm. Theorems considering when MV is optimal for group decisions date back to Condorcet's 1785 jury \emph{decision} theorem. The same error independence assumption underlying the theorem can be used to prove a jury \emph{evaluation} theorem that does purely algebraic evaluation (AE) of juror performance based on a batch of their decisions. Three or more binary jurors are enough to obtain the only two possible statistics of their correctness on a test they took. AE is superior to MV in three ways. First, its empirical assumptions are looser and can handle jurors less than 50\% accurate in making decisions. Second, it has point-like precision in evaluating them given its assumption of error independence. This precision enables a multi-accuracy approach that has higher labeling accuracy than MV and comes with empirical uncertainty bounds. And, third, it is self-alarming about the failure of its error independence assumption. Experiments using demographic data from the American Community Survey confirm the practical utility of AE over MV. Two implications of the theorem for AI safety are discussed - a principled way to terminate infinite monitoring chains (who grades the graders?) and the super-alignment problem (how do we evaluate agents doing tasks we do not understand?).

群体智慧代数评估AI安全评估机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。