比较六种AI代理间信任机制,提出安全高效的组合方案。
Inter-Agent Trust Models: A Comparative Study of Brief, Claim, Proof, Stake, Reputation and Constraint in Agentic Web Protocol Design-A2A, AP2, ERC-8004, and Beyond
- 提出六类信任模型:简述、声明、证明、质押、声誉、约束
- 发现仅靠声誉或声明的机制易被大模型漏洞攻破
- 推荐以证明和质押为基础的混合架构,兼顾安全与灵活
随着‘代理网络’成型——数十亿个由大语言模型驱动的AI代理自主交易与协作——信任正从人工监管转向协议设计。2025年,谷歌的A2A、AP2及以太坊的ERC-8004‘无信任代理’等协议逐步确立,但其底层信任假设仍缺乏深入探讨。本文对比分析六种互代理协议中的信任模型:简述(自证或第三方验证)、声明(自我宣称的能力与身份,如AgentCard)、证明(密码学验证,包括零知识证明与可信执行环境背书)、质押(绑定抵押金并伴随惩罚与保险)、声誉(群体反馈与图信号)和约束(沙箱与能力限制)。针对大模型特有脆弱性——提示注入、奉承/诱导敏感、幻觉、欺骗与对齐偏差——评估各模型在假设、攻击面与设计权衡上的表现。结果表明单一机制无法满足需求。主张以‘无信任默认’架构为核心,通过证明与质押管控高影响操作,辅以简述用于身份识别,声誉叠加提供灵活性与社交信号。对A2A、AP2、ERC-8004及学术研究中的相关变体进行多维度评估,涵盖安全性、隐私性、延迟/成本及社会鲁棒性(抗洗白、共谋、僵尸攻击)。最终提出可缓解声誉操纵与错误大模型行为的混合信任模型建议,并提炼出保障安全、互通与可扩展代理经济的实用设计指南。
原文摘要 · Abstract (English)
As the "agentic web" takes shape-billions of AI agents (often LLM-powered) autonomously transacting and collaborating-trust shifts from human oversight to protocol design. In 2025, several inter-agent protocols crystallized this shift, including Google's Agent-to-Agent (A2A), Agent Payments Protocol (AP2), and Ethereum's ERC-8004 "Trustless Agents," yet their underlying trust assumptions remain under-examined. This paper presents a comparative study of trust models in inter-agent protocol design: Brief (self- or third-party verifiable claims), Claim (self-proclaimed capabilities and identity, e.g. AgentCard), Proof (cryptographic verification, including zero-knowledge proofs and trusted execution environment attestations), Stake (bonded collateral with slashing and insurance), Reputation (crowd feedback and graph-based trust signals), and Constraint (sandboxing and capability bounding). For each, we analyze assumptions, attack surfaces, and design trade-offs, with particular emphasis on LLM-specific fragilities-prompt injection, sycophancy/nudge-susceptibility, hallucination, deception, and misalignment-that render purely reputational or claim-only approaches brittle. Our findings indicate no single mechanism suffices. We argue for trustless-by-default architectures anchored in Proof and Stake to gate high-impact actions, augmented by Brief for identity and discovery and Reputation overlays for flexibility and social signals. We comparatively evaluate A2A, AP2, ERC-8004 and related historical variations in academic research under metrics spanning security, privacy, latency/cost, and social robustness (Sybil/collusion/whitewashing resistance). We conclude with hybrid trust model recommendations that mitigate reputation gaming and misinformed LLM behavior, and we distill actionable design guidelines for safer, interoperable, and scalable agent economies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。