arXiv:2608.00981cs.AI2026-08

提出可验证的双面审计机制,识别AI科学发现中的虚假进步。

Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science, with the Negative Side Decidable

  • 构建可形式化验证的负面审计规则,禁止伪结自由预言机表示交叉碱基对
  • 实测表明,无求解器的算子在43个目标上表现优于基准,但仅2项被独立验证
  • 适用于评估自改进AI系统的真实创新能力,尤其适合审慎验证前沿成果

当自进化AI科学系统声称获得新能力时,证据通常来自基准提升、描述长度阈值或p值。这些均无法区分真实进步、额外搜索、验证器变化或对不可靠预言机的适应。我们建立双面审计机制,其负面对应一个形式事实:无伪结的预言机无法表示交叉碱基对,因此先验验证器范围可在线下精确界定。'新'是相对于智能体自身先前状态,而非基础模型。首先,单个不可靠预言机最多能夸大能力的程度:一个无需求解器的算子在43/60个交叉RNA目标上成功设计,高于上下文无关基线0/60;但在三个预测器下仅1/60存活。在相同43个目标上,该算子从未见过的预测器仅确认2项设计,而最小自由能求解器确认26项(p=8e-7)。系统自身统计量无法察觉此差距。其次,代理编写程序在外部裁判下超越人类编写程序,仅需极少计算资源。六种前沿模型中,两个无超时运行的算子在951组配对单位上以0.293对比我们的0.095([+0.108, +0.297],p=5e-5),且仅消耗4.6–10倍更少的预言机调用。三重阶梯:外部裁定差异已达成,不依赖计算投入(双向达成),机制已识别并可迁移(未达成;测试七种候选均无效)。上限即评审团本身:其三个预测器共享最近邻热力学参数,两两一致性达kappa=0.673。该审计同样严苛审视我们自身系统:匹配的无向搜索为精确零,且无需搜索的探测显示84%的显著效果由随机序列即可达成。

原文摘要 · Abstract (English)

When a self-improving AI-for-science system claims a new capability, the evidence is usually a benchmark delta, a description-length gate, or a p-value. None separates a real gain from extra search, from a changed verifier, or from adaptation to a fallible oracle. We build a two-sided audit whose negative side is a formal fact: a pseudoknot-free oracle provably cannot represent a crossing base pair, so the prior verifier's range is bounded exactly, offline, before any run. "New" is relative to the agent's prior self, never to the base model. First, how far a single fallible oracle can inflate a capability claim. An invented, solver-free operator solves 43/60 crossing RNA targets under the predictor it optimizes, above a context-free floor of 0/60; under three predictors, 1/60 survives. Paired on the same 43 targets, a predictor the operator never saw confirms 2 of its designs against 26 for a minimum-free-energy solver (p = 8e-7). No statistic computed from the system and its own oracle sees that gap. Second, agent-written procedures can beat a human-written one under a judge no objective can flatter, at a fraction of the compute. Of six frontier models, the two whose operators ran without timeouts carry over at 0.293 against our 0.095 (n = 951 paired units, target-clustered [+0.108, +0.297], p = 5e-5) while spending 4.6-10x fewer oracle calls. Three rungs: difference under an outside adjudicator (reached), not bought with compute (reached, both directions), mechanism identified and transferable (not reached; seven candidates tested, none moves the statistic). The ceiling is the panel itself: its three predictors share nearest-neighbour thermodynamic parameters, two agreeing at kappa = 0.673. The audit is as unsparing about our own system: matched undirected search is an exact zero, and a search-free probe puts 84% of our headline effect on targets a random sequence already solves.

AI科研能力审计自进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。