arXiv:2502.04313cs.LGcs.AI2025-02ICML被引 50

大模型越强越相似,导致自动监管易失效。

Great Models Think Alike and this Undermines AI Oversight

  • 用错题重合度衡量模型相似性,发现判别结果偏倚于相似模型。
  • 弱模型指导强模型时,互补知识是性能提升关键。
  • 模型越强越趋同,共性错误增多,监管风险上升。

随着语言模型能力提升,人工评估与监督变得愈发困难。人们期望其他语言模型能自动化完成这些任务,即「AI Oversight」。本文提出一种基于模型错题重合的相似性度量方法——机会调整概率一致率(CAPA),研究模型相似性对监督的影响。实验表明,以模型为裁判的评分偏好与其相似的模型,扩展了近期自偏好研究的结果。在训练阶段,弱监督模型与强学生模型间的互补知识对实现「弱到强泛化」至关重要。随着模型能力增强,其错误更难被发现,我们可能更依赖AI监督。然而,观察到一个令人担忧的趋势:模型错误正随能力提升而趋于相似,暗示存在关联性失败的风险。本工作强调在新兴的AI监督范式中,报告并校正模型相似性的重要性。

原文摘要 · Abstract (English)

As Language Model (LM) capabilities advance, evaluating and supervising them at scale is getting harder for humans. There is hope that other language models can automate both these tasks, which we refer to as ''AI Oversight''. We study how model similarity affects both aspects of AI oversight by proposing Chance Adjusted Probabilistic Agreement (CAPA): a metric for LM similarity based on overlap in model mistakes. Using CAPA, we first show that LLM-as-a-judge scores favor models similar to the judge, generalizing recent self-preference results. Then, we study training on LM annotations, and find complementary knowledge between the weak supervisor and strong student model plays a crucial role in gains from ''weak-to-strong generalization''. As model capabilities increase, it becomes harder to find their mistakes, and we might defer more to AI oversight. However, we observe a concerning trend -- model mistakes are becoming more similar with increasing capabilities, pointing to risks from correlated failures. Our work underscores the importance of reporting and correcting for model similarity, especially in the emerging paradigm of AI oversight.

AI监督模型相似性大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。