arXiv:2507.19486cs.HCcs.AI2025-07被引 2

研究发现简单监督协议易受认知偏见影响,难以可靠验证更强模型。

Confirmation bias: A challenge for scalable oversight

  • 测试人类在已知模型大部分正确时的判断能力
  • 发现协议未提升整体准确性,且研究后信心反而增强
  • 适合关注模型可信赖性与监督机制设计的研究者

可扩展的监督协议旨在让评估者准确验证比自身更强大的AI模型。然而,人类评估者存在认知偏见,可能导致系统性错误。我们开展两项研究,考察评估者在知道模型‘大部分时间正确但非始终正确’情境下的表现。结果表明,所测试的协议整体上无显著优势;在第一项研究中,展示正反论证能提高模型错误时的准确率;第二项研究中,无论答案是否正确,参与者在在线查证后均对系统答案更自信。我们还重新分析了先前研究数据,发现其乐观结论可能源于评估者拥有模型不具备的知识,而这一优势随模型能力提升而减弱。研究强调必须检验监督协议对评估者偏见的鲁棒性,以及其是否优于直接信任模型,且性能能否随问题难度和模型能力增长而持续提升。

原文摘要 · Abstract (English)

Scalable oversight protocols aim to empower evaluators to accurately verify AI models more capable than themselves. However, human evaluators are subject to biases that can lead to systematic errors. We conduct two studies examining the performance of simple oversight protocols where evaluators know that the model is "correct most of the time, but not all of the time". We find no overall advantage for the tested protocols, although in Study 1, showing arguments in favor of both answers improves accuracy in cases where the model is incorrect. In Study 2, participants in both groups become more confident in the system's answers after conducting online research, even when those answers are incorrect. We also reanalyze data from prior work that was more optimistic about simple protocols, finding that human evaluators possessing knowledge absent from models likely contributed to their positive results--an advantage that diminishes as models continue to scale in capability. These findings underscore the importance of testing the degree to which oversight protocols are robust to evaluator biases, whether they outperform simple deference to the model under evaluation, and whether their performance scales with increasing problem difficulty and model capability.

AI监督认知偏见可扩展性评估机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。