arXiv:2601.14262stat.MEcs.AI2026-01

提出评估评估的框架,首次量化比较十种评估方法优劣。

On Meta-Evaluation

  • 构建评估空间与AxiaBench基准,系统对比十种评估方法。
  • 发现现有方法无法兼顾准确与高效,实验设计偏差显著。
  • 新采样方法跨领域表现最优,可提升研究可信度。

评估是实证科学的基础,但对评估本身的元评估仍严重不足。尽管观察研究、实验设计(DoE)和随机对照试验(RCT)塑造了现代科研实践,但它们在不同领域的有效性与实用性缺乏系统比较。本文提出元评估的正式框架,定义评估空间、其结构化表示,并引入名为AxiaBench的基准。该基准首次实现对十种常用评估方法在八个代表性应用领域的规模化定量比较。分析显示:无一现有方法能在多样场景中同时兼顾准确与效率,尤其DoE与观察设计与真实情况存在显著偏差。进一步评估先前元评估研究中的全空间分层采样统一方法,结果表明其在所有测试领域均持续优于现有方法。这些成果确立元评估作为独立科学对象的地位,为计算与实验研究中可信评估的发展提供概念基础与实用工具集。

原文摘要 · Abstract (English)

Evaluation is the foundation of empirical science, yet the evaluation of evaluation itself -- so-called meta-evaluation -- remains strikingly underdeveloped. While methods such as observational studies, design of experiments (DoE), and randomized controlled trials (RCTs) have shaped modern scientific practice, there has been little systematic inquiry into their comparative validity and utility across domains. Here we introduce a formal framework for meta-evaluation by defining the evaluation space, its structured representation, and a benchmark we call AxiaBench. AxiaBench enables the first large-scale, quantitative comparison of ten widely used evaluation methods across eight representative application domains. Our analysis reveals a fundamental limitation: no existing method simultaneously achieves accuracy and efficiency across diverse scenarios, with DoE and observational designs in particular showing significant deviations from real-world ground truth. We further evaluate a unified method of entire-space stratified sampling from previous evaluatology research, and the results report that it consistently outperforms prior approaches across all tested domains. These results establish meta-evaluation as a scientific object in its own right and provide both a conceptual foundation and a pragmatic tool set for advancing trustworthy evaluation in computational and experimental research.

元评估评估框架实验设计可信研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。