HLE多选题测试实则只测单一推理能力,顶尖模型难区分。
Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset

- 用项目反应理论分析428道题,发现仅一个核心能力维度。
- 各领域分数方差仅3.5%由标签解释,顶尖模型得分高度重合。
- 测量精度在强模型所在区域急剧下降,不适合评估前沿模型。
人类最后考试(HLE)广泛用于评估前沿语言模型。其题目分为八个学科领域,子分常被视为不同能力的证据。但尚无研究检验这些标签是否对应可分离的潜在构念,或该基准能否有效区分能力相近的模型。我们对29个大语言模型在HLE纯文本多选题子集(J = 428项)上进行评估,并应用心理测量学方法分析其维度与测量精度分布。拟合双参数逻辑斯蒂克项目反应模型后发现,一致证据表明HLE衡量的仅为单一通用推理因子:McDonald's ω_h = 0.998,领域标签仅解释3.5%的答题方差,领域内与领域间残差相关性几乎相同(Cohen's d = 0.016),领域特定能力估计与总分高度冗余(r ≥ 0.81)。对测验信息函数的独立分析显示,测量精度集中于中等能力水平,在θ=0以上急剧下降,而前沿模型正位于此区间。结果表明,HLE的领域子分不应被解释为独立能力,且其区分最强模型的能力有限。
原文摘要 · Abstract (English)
Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities. However, no study has assessed whether these labels correspond to empirically separable latent constructs, nor whether the benchmark effectively differentiates between models of similar ability. We evaluate 29 LLMs on the text-only multiple-choice subset of HLE ($J = 428$ items) and apply psychometric methods to assess both the dimensionality of the benchmark and the distribution of its measurement precision. Fitting a two-parameter logistic IRT model, we find convergent evidence that HLE measures a single general reasoning factor: McDonald's $ω_h = 0.998$, domain labels explain only 3.5\% of item response variance, within- and between-domain residual correlations are nearly identical (Cohen's $d = 0.016$), and domain-specific ability estimates are near-redundant with the total score ($r \geq 0.81$). A separate analysis of the test information function reveals that measurement precision concentrates at moderate ability levels and drops sharply above $θ= 0$, where frontier models sit. These findings suggest that HLE's domain subscores do not warrant distinct capability interpretations and that the benchmark's ability to discriminate among the strongest models is limited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。