arXiv:2412.08897cs.AIcs.LG2024-12ICLR被引 7

用神经网络设计可验证的交互式证明,让弱验证者也能信任强模型的输出。

Neural Interactive Proofs

  • 构建验证者-证明者博弈框架,统一现有交互协议
  • 在图同构和代码验证任务中验证协议有效性
  • 为可信AI系统提供新思路,适合对安全性要求高的场景

我们研究受信任但计算能力有限的代理(验证者)如何与一个或多个强大但不可信的代理(证明者)交互以完成任务。具体而言,当代理由神经网络表示时,此类解决方案称为神经交互证明。首先,我们提出一个基于证明者-验证者博弈的统一框架,该框架推广了先前提出的交互协议。随后,我们描述了几种生成神经交互证明的新协议,并对新旧方法进行了理论比较。最后,我们在两个领域进行实验以支持该理论:一个展示核心思想的简化图同构问题,以及使用大型语言模型的代码验证任务。我们的目标是为未来神经交互证明的研究及其在构建更安全人工智能系统中的应用奠定基础。

原文摘要 · Abstract (English)

We consider the problem of how a trusted, but computationally bounded agent (a 'verifier') can learn to interact with one or more powerful but untrusted agents ('provers') in order to solve a given task. More specifically, we study the case in which agents are represented using neural networks and refer to solutions of this problem as neural interactive proofs. First we introduce a unifying framework based on prover-verifier games, which generalises previously proposed interaction protocols. We then describe several new protocols for generating neural interactive proofs, and provide a theoretical comparison of both new and existing approaches. Finally, we support this theory with experiments in two domains: a toy graph isomorphism problem that illustrates the key ideas, and a code validation task using large language models. In so doing, we aim to create a foundation for future work on neural interactive proofs and their application in building safer AI systems.

交互证明神经网络可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。