arXiv:2606.05433cs.AIcs.SY2026-06被引 1

让顶尖大模型训练可被可信验证,打破自我申报困局

Zero knowledge verification for frontier AI training is possible

  • 用零知识证明+实时计算承诺,验证真实训练过程
  • 支持浮点计算精确验证,训练开销仅个位数百分比
  • 适合需要可验证治理的前沿AI国际合作场景

前沿AI治理框架日益以累计训练算力作为高影响力模型的主要判定标准,但执行依赖自我报告,因缺乏训练过程的技术验证手段。未来国际协议面临更高风险:历史上的协同监管始终依托技术验证,否则协议仅具声明意义。近期治理分析认为零知识证明是可行方向,但当前在前沿规模下不切实际。本文认为其不切实际源于范式局限而非根本障碍,提出一种面向密集预训练的验证架构,结合预先承诺的训练规范、节点间网络观测与运行中生成的中间计算梅克尔承诺,通过原生支持BF16/FP32的零知识虚拟机(zkVM)进行验证。该方案验证GPU实际执行的浮点运算,而非固定精度近似,并通过私有训练规范保护模型架构机密性。协议生成三类证明:初始化时的创世证明、训练过程中的步骤证明,以及事前承诺的政策相关断言,使训练记录成为可治理的凭证。预计可在约36个月内实现可部署原型,训练侧开销为个位数百分比,远低于需六至十年的专用芯片验证周期。文中列出了13个开放研究与工程问题,构成外部协作的研究议程。

原文摘要 · Abstract (English)

Frontier AI governance frameworks increasingly use cumulative training compute as the primary criterion for designating high-impact models, but enforcement rests on self-reporting because no technical verification primitive for training exists. Any future international agreement on frontier AI faces the same problem at higher stakes: coordinated regulation of technologies with significant externalities has historically rested on technical verification, without which agreements are declaratory. Recent governance analyses judge zero-knowledge proofs a promising candidate but currently impractical at frontier scale [26, 4]. We argue the impracticality is paradigm-bound rather than fundamental, and propose a verification architecture for frontier dense pre-training combining a pre-committed training specification, inter-node network observations, and on-the-fly Merkle commitments of intermediate computation, verified through a zero-knowledge Virtual Machine (zkVM) with native BF16/FP32 precompiles. The proof checks the actual floating-point computation the GPU performed rather than a fixed-point approximation, and preserves model-architecture confidentiality through a private training specification. The protocol produces three proof types: a genesis proof at initialisation, in-training step proofs across the run, and ex-ante attestations enforcing policy-relevant claims as running invariants, turning the training record into a governance-enforceable artefact. We estimate a deployable proof of concept within approximately 36 months at single-digit-percent training-side overhead, against a six-to-ten-year cycle for verification-grade custom silicon. Thirteen open research and engineering problems are catalogued as a research agenda for external contribution

零知识证明模型验证前沿AI治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。