arXiv:2507.00631cs.GTcs.AI2025-07被引 3

用押金博弈机制让不可信的AI代理自动保证任务正确性。

A Protocol for Trustless Verification Under Uncertainty

  • 通过递归验证游戏,让代理以押金承诺任务正确性。
  • 错误方被罚没押金,正确挑战者获奖励,形成激励闭环。
  • 适合构建去中心化AI协作系统的研究者与开发者。

在动态、低信任环境中,自主AI代理需委托子代理完成任务,但无法通过预先设定或集中监管保证正确性。本文提出一种协议,通过抵押声明在递归验证游戏中强制正确性:任务以意图形式发布,求解者竞标执行,完成后由验证者事后检查。任何挑战者可通过押注来质疑结果,触发验证流程。错误代理将被罚没押金,正确挑战者获得奖励,且存在惩罚错误验证者的升级路径。当求解者、挑战者和验证者的激励对齐时,虚假行为将导致自身损失,使正确性成为纳什均衡。

原文摘要 · Abstract (English)

Correctness is an emergent property of systems where exposing error is cheaper than committing it. In dynamic, low-trust environments, autonomous AI agents benefit from delegating work to sub-agents, yet correctness cannot be assured through upfront specification or centralized oversight. We propose a protocol that enforces correctness through collateralized claims in a recursive verification game. Tasks are published as intents, and solvers compete to fulfill them. Selected solvers carry out tasks under risk, with correctness checked post hoc by verifiers. Any challenger can challenge a result by staking against it to trigger the verification process. Incorrect agents are slashed and correct opposition is rewarded, with an escalation path that penalizes erroneous verifiers themselves. When incentives are aligned across solvers, challengers, and verifiers, falsification conditions make correctness the Nash equilibrium.

AI代理博弈机制去中心化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。