arXiv:2608.29463cs.CRcs.AI2026-08

提出五类评测污染类型,帮助识别论文中未披露的模型泄露风险。

Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation

  • 按防御措施分类污染类型:直接、衍生、时间、分布、获取型泄露。
  • 仅保留私有测试集无法完全防范污染,第五类需在评估时记录。
  • 设计可验证的披露协议,支持透明化报告与同行复现。

评测分数是模型、评估框架、激发预算、样本群体及污染状态的联合属性。排行榜只公布模型和分数,导致能力与泄露在观测上等价。现有分类体系用于自动检测污染,而非论文发表时记者面对的核心问题:在已应用缓解措施的前提下,哪些有效性威胁仍存在?本文提出以每类缓解手段所击败的污染类型为组织原则的新分类法,涵盖训练期与评估期的泄漏:直接、衍生、时间、分布与获取型。仅保留私有测试集仅能关闭第一类。第五类即获取型污染发生在评估过程中,因属于单次运行的特性,必须随报告分数一同记录,而非基准发布时。我们将其操作化为包含四个字段的披露协议,允许“未知”作为有效选项,采用CC BY 4.0授权,附带JSON Schema、验证器与实例。两名外部编码员使用预注册工具对41份文档进行评估,变量间线性加权κ值介于0.00至0.35(中位数0.21),在29个主分析文档中,低于注册预期的鲁棒阈值;经机会校正后聚合κ值提升至0.46。两个变量未达预设阈值:分层报告与本文引入的获取型类别。分歧集中于某变量是否适用,而非文档内容本身。仅13%文档报告激发预算,无一文档覆盖全部五类污染。贡献在于提出分类法、伴随的评分侧产物,以及对测量工具可靠性的预注册评估。

原文摘要 · Abstract (English)

A benchmark score is a joint property of the model, the evaluation harness, the elicitation budget, the sampled population, and contamination status. Leaderboards publish the model and the score, so capability and leakage stay observationally equivalent. Existing taxonomies classify contamination for automated detection, not the question a reporter faces at publication: given the mitigations already applied, which validity threats remain open? We introduce a taxonomy organized by the mitigation each type defeats -- direct, derivative, temporal, distributional, and acquired -- spanning training-time and evaluation-time leakage. Holding out a private test set closes the first alone. The fifth is acquired during the evaluation itself; because it is a property of one run, it must be recorded with the reported score rather than with the benchmark release. We operationalize it as a four-field disclosure protocol in which "unknown" is a valid entry, released under CC BY 4.0 with a JSON Schema, a validator, and worked examples. Two coders external to the design team applied a pre-registered instrument to 41 documents. Per-variable linear-weighted $κ$ runs from 0.00 to 0.35 (median 0.21) over 29 main-pass documents against a single-coder test-retest ceiling of 0.84, collapsing under the class skew the registration anticipated; pooling raises it to 0.46 through chance correction rather than better agreement. Two variables fall below the prevalence-robust threshold registered in advance: strata reporting and the acquired type introduced here. Disagreement concentrates on when a variable applies rather than on what a document states. Elicitation budgets are reported in 13% of documents, and no document addresses all five types. The contribution is the taxonomy, the score-side artifact that follows from it, and a pre-registered measurement of instrument reliability and current disclosure.

评测污染方法论透明度可信性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。