构建五个领域错误标注数据集,助力可信AI监督与模型评估
FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research
- 构建含专家标注错误的五类长文本解题数据集
- 发现部分模型在特定任务上表现低于人类专家
- 可用于训练更可靠的模型评判者,支持可扩展监督
随着AI模型处理日益复杂的问题,确保可靠的人类监督变得愈发困难,因验证解决方案难度增加。现有可扩展监督方法包括辩论、批判性评估和证明-验证游戏等。这些方法的有效性依赖于包含(1)长篇专家验证正确解法和(2)带错误标注的有缺陷解法的数据集,但此类数据集极为稀缺。为此,我们提出FindTheFlaws,涵盖医学、数学、科学、编程和洛巴恩语言五个领域的五个多样化数据集。每个数据集包含问题及长篇解法,并附有专家标注以确认正确性或指出具体推理错误。我们评估了前沿模型的批判能力,观察到性能差异:表现较弱的模型可作为更强大模型的裁判/验证者。此外,在某些任务组合中,人类专家基准甚至优于顶尖模型,表明其在可扩展监督实验中更具优势。
原文摘要 · Abstract (English)
As AI models tackle increasingly complex problems, ensuring reliable human oversight becomes more challenging due to the difficulty of verifying solutions. Approaches to scaling AI supervision include debate, in which two agents engage in structured dialogue to help a judge evaluate claims; critique, in which models identify potential flaws in proposed solutions; and prover-verifier games, in which a capable 'prover' model generates solutions that must be verifiable by a less capable 'verifier'. Evaluations of the scalability of these and similar approaches to difficult problems benefit from datasets that include (1) long-form expert-verified correct solutions and (2) long-form flawed solutions with annotations highlighting specific errors, but few are available. To address this gap, we present FindTheFlaws, a group of five diverse datasets spanning medicine, mathematics, science, coding, and the Lojban language. Each dataset contains questions and long-form solutions with expert annotations validating their correctness or identifying specific error(s) in the reasoning. We evaluate frontier models' critiquing capabilities and observe a range of performance that can be leveraged for scalable oversight experiments: models performing more poorly on particular datasets can serve as judges/verifiers for more capable models. Additionally, for some task/dataset combinations, expert baselines exceed even top model performance, making them more beneficial for scalable oversight experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。