arXiv:2510.23191cs.LGstat.ML2025-10被引 12

为机器学习模型评估建立可解释的科学标准,避免误读分数含义。

The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models

  • 基于心理测量学提出评估有效性标准,明确推理前提。
  • 三案例验证:图像识别、天气预测、人生事件可预测性。
  • 适合关注模型评估可信度的研究者和审稿人。

预测性基准测试是机器学习研究中核心的知识实践,常用于科学探究。然而,基准分数仅反映模型在特定数据集和任务上的表现,无法直接推导出关于理论任务(如图像分类)的科学结论。这需要对学习任务结构、评估函数和数据分布做出额外假设。本文借鉴心理测量理论,明确提出构造有效性条件,并通过三个典型案例检验这些假设:用ImageNet衡量计算机视觉工程进展;用WeatherBench评估具有政策意义的天气预测;用Fragile Families Challenge分析人生事件可预测性的局限。该框架明确了基准分数支持各类科学主张的前提条件,将预测性基准测试置于认知论视角下,凸显其作为机器学习中概念与理论推理的关键场域。

原文摘要 · Abstract (English)

Predictive benchmarking, the evaluation of machine learning models based on predictive performance and competitive ranking, is a central epistemic practice in machine learning research and an increasingly prominent method for scientific inquiry. Yet, benchmark scores alone provide at best measurements of model performance relative to an evaluation dataset and a concrete learning problem. Drawing substantial scientific inferences from the results, say about theoretical tasks like image classification, requires additional assumptions about the theoretical structure of the learning problems, evaluation functions, and data distributions. We make these assumptions explicit by developing conditions of construct validity inspired by psychological measurement theory. We examine these assumptions in practice through three case studies, each exemplifying a typical intended inference: measuring engineering progress in computer vision with ImageNet; evaluating policy-relevant weather predictions with WeatherBench; and examining limitations of the predictability of life events with the Fragile Families Challenge. Our framework clarifies the conditions under which benchmark scores can support diverse scientific claims, bringing predictive benchmarking into perspective as an epistemological practice and a key site of conceptual and theoretical reasoning in machine learning.

模型评估认知论基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。