arXiv:2606.03034cs.MAcs.AI2026-06

为智能体网络设计信任层,解决能力自夸与真实能力脱节的问题。

Capability Advertisement as a Market for Lemons: A Trust Layer for Heterogeneous Agent Networks

  • 提出可信度分层架构,用概率描述能力并引入筛选与声誉机制。
  • 证明仅靠声明无法区分可靠与虚假智能体,存在低信任均衡陷阱。
  • 无需重训练模型,可适配现有协议,适合构建可信协作的智能体系统。

大型语言模型(LLM)智能体开始相互委托任务。如模型上下文协议(MCP)和Agent2Agent协议(A2A)允许智能体发布其能力并供他人调用,公开的智能体注册表已出现。但这些协议假设所声明的能力是静态且真实的。实际上,智能体的能力具有概率性、随输入变化、随模型更新漂移,且因自身是语言模型,可能自信地错误描述自己。调用者看到的是智能体声称的能力,而非真实能力,缺乏辨别可靠提供者与擅长伪装者的有效方法。我们指出,这一困境源于‘柠檬市场’问题:当质量隐藏而声明成本低廉时,优质与劣质提供者无法区分,诚实可靠者得不到奖励,市场趋于劣化。经济学提供三种解决方案——信号、筛选和声誉,但当前协议中均未实现。本文贡献包括:(1) 构建故障分类体系,将‘自信错误’归为非对抗性、相关性的拜占庭故障子类,传统容错机制未能覆盖;(2) 建立柠檬市场模型,表明基于信任的协议仅能维持低信任均衡;(3) 提出可信层(Trust Layer),一种轻量、协议无关的窄腰层,位于MCP与A2A之上,添加概率能力描述、筛选机制与声誉系统,在过度宣称的成本超过收益时可实现分离均衡;(4) 推导出委托链的可靠性组合边界,并给出端到端部署论证。该设计无需模型重训练,即使信任锚点缺失或被污染也能渐进退化。

原文摘要 · Abstract (English)

Large language model (LLM) agents have begun to delegate work to one another. Protocols such as the Model Context Protocol (MCP) and the Agent2Agent protocol (A2A) let an agent publish what it can do and let others call it, and public registries of such agents are already appearing. These protocols assume an advertised capability is a static, truthful fact. A real agent is none of these things: its competence is probabilistic, varies with input, drifts when the underlying model is updated, and, because the agent is itself a language model, it can describe itself with complete confidence and be wrong. A caller therefore sees what an agent claims to do, not what it can do, with no principled way to tell a reliable provider from a fluent impostor. We argue these difficulties share one cause: the market for lemons. When quality is hidden and claims are cheap, good and bad providers become indistinguishable, honest reliability goes unrewarded, and the market decays toward its worst participants. Economics offers three remedies, signaling, screening, and reputation, and none are present in today's agent protocols. We make four contributions: (1) a failure taxonomy that names confident-wrong as a non-adversarial, correlated subclass of Byzantine faults that classical fault-tolerance mismodels; (2) a market-for-lemons model showing that faith-based protocols admit only a low-trust equilibrium; (3) the Trust Layer, a thin, protocol-agnostic narrow waist above MCP and A2A that adds probabilistic capability descriptors, screening, and reputation, and admits a separating equilibrium when the cost of sustaining an overclaim exceeds the gain from it; and (4) a reliability-composition bound for delegation chains with an end-to-end placement argument. The design needs no model retraining and degrades gracefully when its trust anchors are absent or corrupt.

智能体信任机制去中心化可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。