arXiv:2608.00794cs.AI2026-08

揭示智能体评估中有效性逐层衰减的系统性问题,提出可量化诊断框架。

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation

  • 构建三层有效性衰减模型,揭示任务生成、人机校准与自动评判的乘积式失效机制。
  • 实证发现82%论文缺失可靠度指标,多基准存在任务设计缺陷和非标准英语偏差。
  • 给出基于心理测量学的八项改进方案,支持不同风险等级的可靠性阈值设定。

智能体评估流程产生的基准分数常被用于决策部署、安全认证和合规声明。目前尚无正式框架描述有效性在各阶段的退化过程。本文提出一个三层乘积型有效性模型:V_total <= V_1 × V_2 × V_3,分别对应任务生成(V_1)、人机模拟器校准(V_2)和自动化判断(V_3)。基于实证估计,若每阶段保留70%有效性,则整体有效性最高仅34%(范围0.22–0.54)。通过55篇已发表评估论文的结构化调研验证该模型,发现约82%的研究采用结构不匹配、不完整或缺失的评分者间信度(IRR)指标,与系统性V_3失效一致。进一步识别出V_1失败(10个主流基准中有7个存在任务有效性缺陷)和V_2校准偏差(模拟器间最大差异达9个百分点,非标准美式英语使用者存在系统性差异)。据此提出八项基于心理测量学的改进建议,并确立分领域可靠性阈值(ICC≥0.70;Cronbach's alpha ≥0.67/0.70/0.80,按后果严重程度分级),为实践者与基准作者提供即时可用的诊断与修正工具。

原文摘要 · Abstract (English)

Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims. No formal framework has yet characterized how validity degrades across the stages of these pipelines. We present a three-layer compounding validity model, V_total <= V_1 x V_2 x V_3, that captures multiplicative degradation across task generation (V_1), human-simulator calibration (V_2), and automated judgment (V_3). Under empirically grounded estimates, a pipeline retaining 70% validity at each stage is at most 34% valid against the intended construct (range 0.22--0.54). We validate the model against a structured survey of 55 published agentic evaluation papers, finding that approximately 82% apply structurally mismatched, incomplete, or absent inter-rater reliability (IRR) metrics---a pattern consistent with systematic V_3 collapse. We further identify empirical evidence of V_1 failures (task validity flaws in 7 of 10 popular benchmarks) and V_2 miscalibration (up to 9 percentage points inter-simulator variance, with systematic demographic disparities for non-Standard American English speakers). We derive eight prescriptions grounded in psychometric science and domain-stratified reliability thresholds (ICC>=0.70; alpha >= 0.67/0.70/0.80 by consequence level) that practitioners and benchmark authors can apply immediately. The framework provides a tractable knowledge-based tool for diagnosing and correcting evaluation pipeline validity before deployment decisions are made.

AI评估有效性可靠性心理测量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。