用对抗机制评测大模型,即使人类无法理解任务也能判断对错。
Benchmarking at the Edge of Comprehension
- 设计对抗性评测框架,让人类只验证局部命题
- 8个前沿大模型在数学任务上得分稳定且与外部指标相关
- 适合评估人类难以理解的复杂任务,尤其适用于超大规模模型
随着前沿大语言模型在新基准发布后迅速达到饱和,传统评测面临挑战:若模型持续进步,人类将难以生成有区分性的题目、提供准确答案或评估复杂解法。我们称此为后理解时代。本文提出批判鲁棒评测(Critique-Resilient Benchmarking),一种在人类无法全面理解任务时仍可比较模型的对抗性框架。其核心是批判鲁棒正确性:只要无对手能有力反驳,答案即视为正确。人类作为有限验证者,仅需关注局部断言,从而在超越完全理解的前提下保持评测完整性。采用分项双分图布拉德利-特里模型,联合排名模型解题能力与生成难题的能力。我们在八个前沿大模型的数学领域验证了该方法的有效性,结果显示评分稳定,并与外部能力度量高度相关。本框架将评测重构为人类作为最终裁决者的对抗生成-评估游戏。
原文摘要 · Abstract (English)
As frontier Large Language Models (LLMs) increasingly saturate new benchmarks shortly after they are published, benchmarking itself is at a juncture: if frontier models keep improving, it will become increasingly hard for humans to generate discriminative tasks, provide accurate ground-truth answers, or evaluate complex solutions. If benchmarking becomes infeasible, our ability to measure any progress in AI is at stake. We refer to this scenario as the post-comprehension regime. In this work, we propose Critique-Resilient Benchmarking, an adversarial framework designed to compare models even when full human understanding is infeasible. Our technique relies on the notion of critique-resilient correctness: an answer is deemed correct if no adversary has convincingly proved otherwise. Unlike standard benchmarking, humans serve as bounded verifiers and focus on localized claims, which preserves evaluation integrity beyond full comprehension of the task. Using an itemized bipartite Bradley-Terry model, we jointly rank LLMs by their ability to solve challenging tasks and to generate difficult yet solvable questions. We showcase the effectiveness of our method in the mathematical domain across eight frontier LLMs, showing that the resulting scores are stable and correlate with external capability measures. Our framework reformulates benchmarking as an adversarial generation-evaluation game in which humans serve as final adjudicators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。