arXiv:2608.21496cs.LGmath.ST2026-08

揭示AI验证中最佳N选一的精确可信度边界,解决采样噪声与结构盲区的权衡问题。

The geometry of AI validation: Exact certification limits for iid best-of-N search

  • 基于独立同分布的最佳N选一搜索,建立可靠性表面的核函数模型。
  • 在特定条件下,可靠度已知时,模糊宽度精确为 $B_{m,N}=1+2 extstyleigsum_{r=1}^{m}(-1)^r ext{cos}^{2N}{rπ/[2(m+1)]}$。
  • 提出双门审计规则:先覆盖结构方向,再加独立任务提升精度,适合高可靠性部署场景。

AI系统越来越多地生成多种备选方案、审查证据并部署选定输出。验证因此具有目标相关性:证据仅在干预所解决的方向上支持部署。我们将验证与部署规则表示为可靠性曲面上的核函数。其跨度几何将减少采样噪声的复制与减少结构性盲区的新干预方向区分开来。我们对独立同分布的最佳N选一搜索使该原则精确化。在标量排序、随机平局、最大选择、有界二元真理及稳定秩-真值关系下,通过 $n=m$ 知道最佳N选一的可靠性,可得精确模糊宽度 $B_{m,N}=1+2\sum_{r=1}^{m}(-1)^r\cos^{2N}{rπ/[2(m+1)]}$。在完全有界世界中可达整个区间,且完整前缀在 $n\le m$ 的可靠性均值审计中信息最大化。主导尺度为 $m^2/N$:当 $m$ 与 $\sqrt{N}$ 同阶时,模糊度仍约为0.83;而宽度 $\varepsilon$ 要求 $m$ 为 $\sqrt{N\log(1/\varepsilon)}$ 量级。单调性给出精确一致逼近前沿;Lipschitz界导出精确截尾对偶和阶尖锐的 $L/m^2$ 模糊度。这些结果催生双门审计规则:先确立结构覆盖,再添加独立任务以提高精度。对数学推理与代码选择的回顾性研究构建了兼容的部署值,显示显著分离,并表明在82个发现任务上冻结的得分尾部审计规则能显著降低保留误差。超越独立同分布搜索,该几何仅适用于已知或独立估计的核函数;实证分析具说明性,非前瞻性干预。

原文摘要 · Abstract (English)

AI systems increasingly generate alternatives, inspect evidence, and deploy a selected output. Validation is therefore target-relative: evidence certifies deployment only in directions resolved by the interventions that produced it. We represent validation and deployment rules as kernels over a reliability surface. Their span geometry separates replication, which reduces sampling noise, from new intervention directions, which reduce structural blindness. We make this principle exact for iid best-of-$N$ search. Under scalar ranking, randomized ties, maximum selection, bounded binary truth, and a stable rank-truth relation, knowing best-of-$n$ reliability through $n=m$ leaves exact ambiguity width $B_{m,N}=1+2\sum_{r=1}^{m}(-1)^r\cos^{2N}{rπ/[2(m+1)]}$. Explicit bounded worlds attain the entire interval, and the complete prefix is information-maximal among reliability-mean audits confined to $n\le m$. The governing scale is $m^2/N$: when $m$ is proportional to $\sqrt{N}$, ambiguity remains about 0.83, while width $\varepsilon$ requires $m$ of order $\sqrt{N\log(1/\varepsilon)}$. Monotonicity gives an exact uniform-approximation frontier; a Lipschitz bound gives an exact capped-tail dual and order-sharp $L/m^2$ ambiguity. These results yield a two-gate audit rule: establish structural coverage, then add independent tasks for precision. Retrospective studies of mathematical reasoning and code selection construct compatible deployment values with wide separation and show that a score-tail audit rule frozen on 82 discovery tasks substantially reduces held-out error. Beyond iid search, the geometry applies only to known or independently estimated kernels; the empirical analyses are illustrative rather than prospective interventions.

AI验证最佳选择可信度边界统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。