无需对战,单模型也能安全验证AI输出。
How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs
- 用单个模型替代辩论机制,通过交互式证明验证AI输出
- 在存在错误查询的噪声环境下仍能保持结果稳定
- 适合需要高可靠性的AI系统安全验证场景
随着人工智能模型能力不断增强,确保其输出与人类意图一致变得至关重要。现有研究多采用辩论机制,即两个能力相当的模型相互辩驳以说服弱验证者(如人类)。但该机制假设双方能力对等且至少一方诚实,现实未必成立。本文提出无需辩论的新方案:研究单证明者交互式证明,用于AI安全验证。此前单证明者方法不适用于含外部查询(如人工判断或网络数据库)的计算场景。本文提出双重高效单证明者交互式证明与论证,适用于带预言机访问的计算,在以下两种情形中有效:(1) 计算具有鲁棒性,即最多少量预言机查询出错时输出不变;(2) 预言机为低次多项式。结果表明,即使无辩论,只要预言机结构良好或抗噪声,交互验证仍可实现。
原文摘要 · Abstract (English)
As AI models continue to develop powerful capabilities, it becomes critical that we are able to verify that their output is aligned with our intentions. A recent line of work focuses on verification via debate, a model of interactive proofs where two competing powerful provers, or AI models, debate each other to convince a weak verifier, or a human, of the correctness of their claim. However, debate assumes that the two AI models possess equal abilities and that one of them is truthful, which may not be realistic. In this work, we show \emph{how to avoid debate}: we initiate the study of \emph{single-prover} interactive proofs for AI safety. Prior results in single-prover interactive proofs do not immediately carry over to the AI safety setting: for example, they do not work when the computation has access to an oracle, such as to human judgment or an external database such as the web. We present doubly-efficient single-prover interactive proofs and arguments for oracle-aided computations (also known as relativizing proofs), in the settings where (1) the computation is robust, in the sense that the output does not change if at most a small fraction of the answers to oracle queries are incorrect, or (2) the oracle is a low-degree polynomial. These results suggest that interactive verification is possible even without debate, under structured or noise-tolerant oracle access.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。