arXiv:2605.14167cs.AIcs.CY2026-05被引 1

AI基准测试隐含理论假设,可能固化范式局限。

The Evaluation Trap: Benchmark Design as Theoretical Commitment

  • 提出'认知学'方法,从能力主张推导评估标准
  • 发现现有基准常无法区分真实能力与替代行为
  • 适合反思评估体系的研究者与模型设计者

每个AI基准都内嵌了关于能力的理论假设。当这些假设未经审视而成为固定信念时,基准会通过限定进步标准来巩固主导范式。长期来看,狭隘的评估重构了能力概念:架构与定义被选择以适应基准可读性,导致评估不再追踪独立目标,反而生成由自身操作假设定义的版本。结果形成评估陷阱:自强化的评估被视为有效,既制造又遮蔽了当前范式的能力边界。本文提出‘认知学’方法,从技术能力主张直接推导评估标准,并审计基准是否能区分真实能力与代理行为。贡献为元评估:一套审计流程、失败模式分类和评估一致性设计准则。通过分析Dupoux等(2026)的提案,揭示其虽在架构层面挑战主流假设,却在评估标准中复制了原有约束,使该限制无法被评估检测而进一步固化。

原文摘要 · Abstract (English)

Every AI benchmark operationalizes theoretical assumptions about the capability it claims to assess. When assumptions function as unexamined commitments, benchmarks stabilize the dominant paradigm by narrowing what counts as progress. Over time, narrow evaluation reorganizes capability concepts: architectures and definitions are selected for benchmark legibility until evaluation ceases to track an independent object and instead produces a version of the target defined by its own operational assumptions. The result is a trap: evaluation frameworks treat self-reinforcing assessments as valid, both creating and obscuring structural limits on what the current paradigm can accomplish. We introduce Epistematics, a methodology for deriving evaluation criteria directly from technical capability claims and auditing whether proposed benchmarks can discriminate the claimed capability from proxy behaviors. The contribution is meta-evaluative: an audit procedure, a failure mode taxonomy, and benchmark-design criteria for evaluating capability-evaluation coherence. We demonstrate the procedure through a worked audit of Dupoux et al. (2026), a proposal that revises the dominant paradigm's theoretical assumptions at the architectural level while reproducing them in its evaluation criteria, thereby entrenching the constraint it seeks to overcome in a form the evaluation cannot detect.

评估陷阱基准设计认知学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。