检验了人类与大模型在教育测评中是否测量相同能力。
Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis

- 用因子分析比较人类与六种多模态大模型的答题结构
- 发现两类群体在化学和数理题中因子结构系统性不同
- 质疑用传统考试评估AI能力的有效性,适合研究评测方法者
大语言模型(LLMs)的快速发展使得评估其能力变得愈发重要。当前普遍做法是使用原本为人类设计的标准化测评工具(如高中化学考试、大学入学考试中的量化推理部分)来评估大模型,并据此推断其具备与人类相似的通用能力。然而,这种推断需满足一个前提:测评工具所揭示的潜在能力结构在人类与大模型之间应具有一致性。本研究通过案例分析,对比了六种多模态大模型与真实人类在上述两个教育场景下的答题数据,采用探索性因子分析、因子一致性及重抽样方法,评估其潜在结构的相似性。结果表明,在两项测评中,人类与大模型的因子结构存在系统性差异,说明这些测评可能并未在两类主体间测量相同的能力。这一发现挑战了以教育测评作为大模型能力推断依据的有效性。
原文摘要 · Abstract (English)
The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities. A common approach is to evaluate LLMs using assessment instruments originally designed to measure skills and competencies in humans, such as standardized exams, and to use performance on these instruments as evidence for generalizable claims about LLMs' underlying abilities on the same skills the assessments are intended to measure in humans. However, from a validity perspective, such inferences require that the relationship between observed performance and underlying constructs established for humans also holds for LLMs. In particular, a necessary condition for transferring score interpretations is similarity in the latent structure of responses to the assessment. In this study, we examine whether this condition holds in two educational contexts: high-school chemistry and a quantitative reasoning section of a university entrance exam. Using a case study design, we compare human response data with responses generated by six multimodal LLMs. Our analytical approach combines exploratory factor analysis, factor congruence, and resampling to assess latent structure similarity across human learners and LLMs. Across both instruments, we find systematic differences between human and LLM factor structures, showing evidence that the analyzed assessments may not measure the same constructs for humans and LLMs. These findings call into question the validity of evaluation practices that use educational assessments to make claims about AI capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。