arXiv:2607.26191cs.AIcs.CL2026-07被引 1

评估分数会过期,应像证据一样标注可信度和时效性。

Position: Evaluation Scores Are Perishable Knowledge Claims

论文配图:Position: Evaluation Scores Are Perishable Knowledge Claims
图 1 · 摘自论文原文
  • 用最弱信号决定整体评分,避免高估模型性能。
  • 均值聚合让排名错乱,前五名模型在两种方法下完全不同。
  • 建议给评估结果加元数据:可信等级、适用范围、失效时间。

语言模型的评估方法越来越依赖多源信号,包括自动指标、大模型评分、人工评估和基准测试结果。当这些信号通过平均聚合时,评估信心可能远超最弱信号的可靠性,这种现象称为评估中的信任膨胀。我们主张将评估分数视为具有三个属性的知识声明:正式性(人工评估比自动指标更强)、作用范围(基准结果仅适用于特定分布,非普适)和有效窗口(随污染积累和分布变化而过期)。多个研究方向(链式思维分析、可能性逻辑、代数理论)支持以最弱环节聚合作为保守终点,其由单一悲观参数控制。基于此,并结合构建智能体评估框架的经验,我们提出评估结果应携带明确元数据(正式性层级、作用范围声明、失效日期),以透明化其认知状态。我们在公开的HELM排行榜上验证了均值聚合的代价:在54个前沿模型、10种场景中,按均值得分和最弱环节得分排序的前五名模型完全不重叠。

原文摘要 · Abstract (English)

Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results. When these signals are aggregated via averaging, evaluation confidence can then substantially exceed the reliability of the weakest signal: a phenomenon we call trust inflation in evaluation. We argue that evaluation scores should be treated as epistemic claims with three properties: formality (human evaluation provides stronger evidence than an automated metric), scope (a benchmark result applies to the tested distribution, not universally), and validity windows (benchmark results expire as contamination accumulates and distributions shift). Several converging research traditions (chain-of-thought analysis, possibilistic logic, and algebraic theory) establish weakest-link aggregation as the conservative endpoint of a parameterized operator family controlled by a single pessimism parameter. Drawing on those traditions, and on concrete lessons from building an evaluation harness for agentic AI, we propose that evaluation results carry explicit metadata (formality tier, scope declaration, and expiration date) to make their epistemic status transparent. We illustrate the cost of mean aggregation on the public HELM leaderboard: across 54 frontier models on ten scenarios, the top-five models ranked by mean score and by weakest-link are completely disjoint.

评估方法可信度模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。