arXiv:2601.18278cs.LGcs.AI2026-01

模型输出被当作测量值时,预测好不等于测得准。

What Do Learned Models Measure?

  • 将模型视为测量工具,关注其输出在不同情况下的稳定性
  • 相同预测性能的模型可能产生系统性不同的测量结果
  • 适合关注模型可靠性与可解释性的研究者

在众多科学与数据驱动的应用中,机器学习模型正越来越多地被用作测量仪器,而非仅预测预设标签。当测量函数由数据学习得到时,观测值到量值的映射由训练分布和归纳偏置隐式决定,导致多个非等价映射均可满足标准预测评估标准。我们形式化了学习型测量函数作为独立的评估焦点,并引入测量稳定性这一属性,捕捉量值在可接受的学习过程实现与不同情境间的不变性。研究表明,机器学习中的标准评估指标,包括泛化误差、校准度和鲁棒性,并不能保证测量稳定性。通过一个真实案例研究,我们展示了具有相似预测性能的模型可能实现系统性不等价的测量函数,分布偏移提供了这一失败的明确例证。综上,我们的结果凸显了现有评估框架在学习模型输出被视为测量值场景下的局限性,提示需要增加新的评估维度。

原文摘要 · Abstract (English)

In many scientific and data-driven applications, machine learning models are increasingly used as measurement instruments, rather than merely as predictors of predefined labels. When the measurement function is learned from data, the mapping from observations to quantities is determined implicitly by the training distribution and inductive biases, allowing multiple inequivalent mappings to satisfy standard predictive evaluation criteria. We formalize learned measurement functions as a distinct focus of evaluation and introduce measurement stability, a property capturing invariance of the measured quantity across admissible realizations of the learning process and across contexts. We show that standard evaluation criteria in machine learning, including generalization error, calibration, and robustness, do not guarantee measurement stability. Through a real-world case study, we show that models with comparable predictive performance can implement systematically inequivalent measurement functions, with distribution shift providing a concrete illustration of this failure. Taken together, our results highlight a limitation of existing evaluation frameworks in settings where learned model outputs are identified as measurements, motivating the need for an additional evaluative dimension.

模型评估测量理论稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。