arXiv:2505.10573cs.CYcs.LG2025-05被引 61

构建评估框架,让AI能力测试更可信。

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

  • 基于心理测量学拆解评估有效性,区分测试结果的真实含义。
  • 证明数学测试表现未必代表通用推理能力,避免夸大结论。
  • 适合研究者、评测人员和政策制定者用于严谨评估AI能力。

尽管人工智能系统的能力与实用性不断进步,但其评估标准却滞后。许多宏大宣称(如模型具备通用推理能力)仅依赖狭窄基准测试(如研究生水平考试成绩),导致评估结果片面且易误导。本文提出一个以有效性为核心的结构化评估框架,帮助判断某项测试表现是反映特定任务掌握程度,还是体现更广泛的推理能力。该框架适用于当前机器学习中多方提供测量数据、下游用户据此验证主张的生态。通过借鉴心理测量学对有效性的分层分析,评估可聚焦关键维度,提升实证效用与决策质量。文中通过视觉与语言模型的案例研究,展示明确考虑有效性如何强化评估证据与主张之间的关联性。

原文摘要 · Abstract (English)

While the capabilities and utility of AI systems have advanced, rigorous norms for evaluating these systems have lagged. Grand claims, such as models achieving general reasoning capabilities, are supported with model performance on narrow benchmarks, like performance on graduate-level exam questions, which provide a limited and potentially misleading assessment. We provide a structured approach for reasoning about the types of evaluative claims that can be made given the available evidence. For instance, our framework helps determine whether performance on a mathematical benchmark is an indication of the ability to solve problems on math tests or instead indicates a broader ability to reason. Our framework is well-suited for the contemporary paradigm in machine learning, where various stakeholders provide measurements and evaluations that downstream users use to validate their claims and decisions. At the same time, our framework also informs the construction of evaluations designed to speak to the validity of the relevant claims. By leveraging psychometrics' breakdown of validity, evaluations can prioritize the most critical facets for a given claim, improving empirical utility and decision-making efficacy. We illustrate our framework through detailed case studies of vision and language model evaluations, highlighting how explicitly considering validity strengthens the connection between evaluation evidence and the claims being made.

AI评估有效性心理测量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。