arXiv:2605.07002cs.AImath.ST2026-05

为生成式AI的自适应审计设计了可随时验证的统计方法,提升评估效率与可靠性。

Adaptive auditing of AI systems with anytime-valid guarantees

论文配图:Adaptive auditing of AI systems with anytime-valid guarantees
图 1 · 摘自论文原文
  • 通过双视角假设检验,将审计过程建模为‘下注式测试’,实现动态采样下的严格统计推断。
  • 在仅20个样本的情况下仍能保持严格第一类错误控制,显著优于固定规则的测试方法。
  • 适合需要快速、可信评估AI系统安全性的研究者与工程师,尤其适用于小样本场景。

生成式AI系统故障模式分析的一大瓶颈在于标注与评估的成本和耗时。因此,自适应测试范式逐渐流行,即根据已有结果动态决定采样案例和数量。然而,这种高度灵活性破坏了经典统计学的假设:观测数通常受限(常为10至50例),且采样与停止决策在数据收集过程中实时做出,而非预先设定。为厘清此类高度自适应审计可得出何种统计结论,本文引入一种双视角假设检验框架:(i) 模型零假设——不存在性能低于目标阈值的故障模式;(ii) 审计员零假设——其采样策略足以发现故障模式。基于安全任意时间有效推断(SAVI),我们将审计建模为‘测试下注’,转化为对两个对立零假设的同时e过程。进一步证明,若审计员足够强大,这两个假设在渐近意义上互为逆命题,即通过严格审计确实可证明AI系统全局鲁棒。实验表明,所提方法在任意时间点均保持严格的类型一误差控制,优于预设规则测试,并可在仅20个观察值时得出统计严谨结论。

原文摘要 · Abstract (English)

A major bottleneck in characterizing the failure modes of generative AI systems is the cost and time of annotation and evaluation. Consequently, adaptive testing paradigms have gained popularity, where one opportunistically decides which cases and how many to annotate based on past results. While this framework is highly practical, its extreme flexibility makes it difficult to draw statistically rigorous conclusions, as it violates classical assumptions: the number of observations is typically limited (often 10 to 50 cases) and decisions regarding sampling and stopping are made in the midst of data collection rather than based a pre-specified rule. To characterize what statistical inferences can be drawn from highly adaptive audits, we introduce a hypothesis testing framework from two 'dueling' perspectives: (i) the model's null that asserts there is no failure mode with performance below a target threshold versus (ii) the auditor's null that asserts they have a sampling strategy that will uncover a failure mode. Leveraging Safe Anytime-Valid Inference (SAVI), we formalize the auditor as conducting 'testing by betting', which translates into simultaneous e-processes for testing the dueling null hypotheses. Furthermore, if the auditor is sufficiently powerful, we prove that these two hypotheses are asymptotically inverses of each other, in that passage of a stringent audit does in fact certify the AI system as being globally robust. Empirically, we demonstrate that our proposed testing procedures maintain anytime-valid type-I error control, outperform pre-specified testing methods, and can reach statistically rigorous conclusions sometimes with as few as 20 observations.

AI审计统计推断自适应测试任意时间有效性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。