arXiv:2606.11417cs.LGcs.AI2026-06

用密封审计的符号压缩进度,可防止代理伪造学习成果。

Signed Compression Progress on a Sealed Audit is Goodhart-Resistant

  • 以密封审计的符号压缩损失变化作为内在奖励
  • 累计奖励精确等于审计性能提升,防作弊
  • 适合关注模型真实进步的可信评估场景

压缩进度是内在动机的经典设想:当世界模型对经验的预测或压缩能力提升时给予奖励。我们证明其可信性:若内在奖励为固定密封审计损失的符号下降,即 r_t = E(theta_{t-1}) - E(theta_t),则累计奖励恰好等于终点审计改进。在有限审计样本下,累计经验奖励最多超出真实改进 2 Delta_n(F, delta),即模型类的均匀审计偏差。该结果与时间无关,一旦密封面板控制住模型类,自适应无需额外代价。但若进度被截断、基于代理自身数据流评分、暴露于高容量模型或使用使 Delta_n 无意义的神经网络类,则保证失效。我们用 Lean 4 机械化了核心结构(伸缩性、有限审计界、有限吉布斯、熵下限),并在 ARC-TGI 网格变换生成器上进行了自适应留出攻击实验。实验验证理论:有限审计偏差随样本数呈 n^{-0.527} 衰减;符号进度能抵抗截断农耕、数据流泄露和噪声电视好奇心;朴素可复用审计易被黑箱标量反馈攻破,而标准发布防御可将攻击控制在 2 Delta_n 以下。密封审计上的符号压缩进度是真实改进的会计信号。

原文摘要 · Abstract (English)

Compression progress is a long-standing proposal for intrinsic motivation: reward an agent when its world model becomes better at predicting or compressing experience. The folk claim is that this reward is "credible" because it is paid only for learning. We make this precise and prove it. If intrinsic reward is the signed decrease of a fixed sealed-audit loss, r_t = E(theta_{t-1}) - E(theta_t), then cumulative reward telescopes exactly to endpoint audit improvement, so no policy can push reward up indefinitely while true audit performance stagnates or degrades. For finite audit panels the same result holds with a sharp false-positive budget: cumulative empirical reward is at most true audit improvement plus 2 Delta_n(F, delta), the uniform audit deviation of the model class. This is horizon-free: adaptivity over time costs nothing once the sealed panel uniformly controls the class. The theorem also identifies the failure modes: the guarantee disappears if progress is clipped, scored on the agent's own stream, exposed to a high-capacity model on a reusable panel, or applied to a neural class that makes Delta_n vacuous. We give a Lean 4 mechanization of the structural core (telescoping, the finite-audit bound, finite Gibbs, and the entropy floor) and an experiment suite on ARC-TGI grid-transformation generators with adaptive holdout attacks. Experiments confirm the theory: finite-audit deviation scales as n^{-0.527}; signed progress resists clip-farming, stream leakage, and noisy-TV curiosity; naive reusable audits are exploitable by black-box scalar feedback, while standard release defenses keep the attack below the 2 Delta_n threshold. Signed compression progress on a sealed audit is an accounting signal of genuine improvement.

内在动机可信评估审计安全压缩进度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。