用模型自动生成挑战题,突破人类能力极限的智能评估新方法
Measuring Intelligence Beyond Human Scale
- 让模型生成公开挑战题,通过竞争筛选其他系统
- 构建可随智能水平扩展的对抗性心理测量评分体系
- 适合评估超越人类的AI系统,无需人工裁判
如何衡量超越人类能力的智能?当前人类设计的评测基准已趋于饱和,而超出人类能力范围的任务,评价者可能无法判断其难度与可验证性。我们指出,绝对尺度评估存在固有局限,提出基于相对测量的新范式:让模型生成公开挑战,以区分其他系统的能力。聚合这些结果形成对抗性心理测量评分系统,可随被测系统的进化而扩展。我们设计了实用协议,降低私有信息攻击动机,实现无需裁判的自动判别,并在可验证与开放域任务中验证该框架,证明模型自生成评估能持续衡量超越人类智能的系统。
原文摘要 · Abstract (English)
How can we measure intelligence beyond human capability? Human-authored benchmarks saturate, and above human capability, examiners may not know which tasks are both hard and verifiable. We argue that this difficulty is inherent to absolute-scale evaluation and propose a new paradigm based on relative measurement in which models generate public challenges that separate other systems. Aggregating these outcomes yields an adversarial psychometric rating system that can scale with the systems being measured. We describe practical protocols that reduce incentives for private-information attacks, support judge-free adjudication, and naturally scale with agent capabilities. We instantiate the framework across verifiable and open-ended, non-verifiable domains, illustrating how model-generated evaluation can continue to measure systems beyond the human frontier.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。