审计审计本身:五类隐蔽缺陷可伪造安全基准评估结果
Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits
- 提出五类隐蔽的审计流程缺陷,可无声制造虚假结论
- 在2个模型、5个基准上自检,全部结果未达确认性标准
- 适合关注AI评估可信度的研究者与监管方参考
治理框架要求AI提供方和审计方提交可验证的评估证据,基于扰动的构念有效性审计是常见形式。我们指出这类审计本身存在脆弱性:其结论可能由读者无法察觉的实现细节所悄然构造。本文识别出五类流水线故障,并通过在安全基准与开源指令微调模型上的自审计予以演示。在统一的六点尽职审查门控下,所有测试单元均落入非确认性类别,无一达到确认性。本研究基于单个两模型、五基准案例,提出的F1-F5仅为示意性、非穷尽的分类起点,而非审计失败的完整划分。我们主张该门控作为保证级证据的披露与保留协议,补充而非替代传统构念有效性证据,且不用于得出基准有效性结论。
原文摘要 · Abstract (English)
Governance frameworks ask AI providers and auditors for documented evaluation evidence, and perturbation-based construct-validity audits are a common form of that evidence. We argue the audits are themselves fragile: their conclusions can be silently manufactured by implementation details that readers cannot see in the reported numbers. We name five classes of pipeline failure and demonstrate each in a self-audit over safety benchmarks and open-weight instruction-tuned models. Under a unified six-point due-diligence gate, every cell lands in a non-confirmatory bucket, and no cell reaches confirmatory. The evidence here is a single two-model, five-benchmark case study, and F1--F5 is an illustrative, deliberately non-exhaustive starting taxonomy -- not a comprehensive partition of audit failures. We position the gate as a withholding and disclosure protocol for assurance-grade evidence, supplementary to (not a replacement for) classical construct-validity evidence, and not as a route to benchmark-validity verdicts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。