用全球人口标准重估AI能力,让评测更真实可信。
From Human-Level AI Tales to AI Leveling Human Scales
- 构建基于全球人口的多能力分级体系,每级对应成功率对数分布。
- 通过教育与认知基准数据估算尺度基底,实现跨任务统一评测。
- 适合关注AI评估公平性、跨文化通用能力的研究者。
将AI模型与“人类水平”对比常因基准不一致或人类基线样本过窄而产生误导。为此,我们提出一个以“全球人口”为校准基准的框架,报告在统一人类锚定尺度上的性能。具体而言,针对推理、理解、知识、体量等不同能力,构建多层级量表,每一级代表全球人口在该任务上成功概率的对数刻度(底数为 $B$)。通过整合涵盖教育与推理的公开人类测试数据(如PISA、TIMSS、ICAR、UKBioBank、ReliabilityBench),利用大语言模型对两类人口特征样本进行外推,估计尺度基底 $B$。通过分组切片与后分层法评估映射质量,新方法实现了相对于全人类群体的量表再校准与标准化。
原文摘要 · Abstract (English)
Comparing AI models to "human level" is often misleading when benchmark scores are incommensurate or human baselines are drawn from a narrow population. To address this, we propose a framework that calibrates items against the 'world population' and report performance on a common, human-anchored scale. Concretely, we build on a set of multi-level scales for different capabilities where each level should represent a probability of success of the whole world population on a logarithmic scale with a base $B$. We calibrate each scale for each capability (reasoning, comprehension, knowledge, volume, etc.) by compiling publicly released human test data spanning education and reasoning benchmarks (PISA, TIMSS, ICAR, UKBioBank, and ReliabilityBench). The base $B$ is estimated by extrapolating between samples with two demographic profiles using LLMs, with the hypothesis that they condense rich information about human populations. We evaluate the quality of different mappings using group slicing and post-stratification. The new techniques allow for the recalibration and standardization of scales relative to the whole-world population.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。