arXiv:2509.19590cs.AIcs.CY2025-09被引 2

AI评估需基于明确能力理论,否则结果可能失真。

Position: AI Evaluations Should be Grounded on a Theory of Capability

  • 将评估视为基于能力理论的推断任务,而非直接测量。
  • 实验证明评估结果受建模假设影响显著。
  • 适合关注评估可信度的研究者与审稿人。

生成模型的评估如今无处不在,其结果深刻影响公众与科学界对AI能力的认知。然而对其可靠性的质疑日益加剧:报告的准确率是否真实反映模型能力?尽管基准结果常被视为能力的直接度量,实际上它们是推断——将得分视为能力证据,已预设了关于任务能力的理论。本文主张AI评估应作为基于明确能力理论的推断任务。虽然这一视角在心理测量学中已是标准做法,但在AI评估中仍发展不足,核心假设常被隐去。以概念验证为例,我们实证表明报告性能强烈依赖评估者的建模假设,凸显透明、理论驱动评估实践的必要性。最后,我们提出评估卡片(Evaluation Card),帮助研究者记录、辩护并审查评估背后的建模决策。

原文摘要 · Abstract (English)

Evaluations of generative models are now ubiquitous, and their outcomes critically shape public and scientific expectations of AI's capabilities. Yet skepticism about their reliability continues to grow. How can we know that a reported accuracy genuinely reflects a model's underlying performance? Although benchmark results are often presented as direct measurements of capability, in practice they are inferences: treating a score as evidence of capability already presupposes a theory of what it means to be capable at a task. We argue that AI evaluations should instead be framed as inference tasks grounded on an explicit theory of capability. While this perspective is standard in fields like psychometrics, it remains underdeveloped in AI evaluation, where core assumptions are often left implicit. As a proof-of-concept, we empirically show that reported performance can depend strongly on the evaluator's modeling assumptions, underscoring the need for transparent, theory-driven evaluation practices. We conclude by offering an Evaluation Card to help researchers document, justify, and scrutinize the modeling decisions underlying AI evaluations.

AI评估能力理论可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。