构建可信代理网络需从架构设计入手,而非事后修补。
Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On

- 提出四支柱框架,从源头设计代理间信任机制。
- 指出现有对齐技术无法应对代理协作中的系统性风险。
- 适合关注智能体协同安全与可信系统的研究人员。
大型语言模型的快速发展催生了具备复杂推理与执行能力的自主型LLM代理。当这些代理从孤立运作转向协作生态时,代理间(A2A)网络应运而生,即异构代理自主协调完成多步骤任务。尽管此类网络相比单一代理完成全部任务更具性能优势,但引入了对抗性组合、语义错位和级联操作失败等系统性漏洞,现有代理对齐技术难以应对。本文认为,A2A网络的可信性无法通过针对单个代理设计的现有协议进行事后补救,而必须在协作框架之初就嵌入信任机制。为此,我们提出一个综合性概念框架,通过四个设计支柱确立A2A系统的信任基础。
原文摘要 · Abstract (English)
The rapid advancement of Large Language Models has given rise to autonomous LLM-based agents capable of complex reasoning and execution. As these agents transition from isolated operation to collaborative ecosystems, we witness the emergence of the Agent-to-Agent (A2A) network, a paradigm where heterogeneous agents autonomously coordinate to solve multi-step tasks. While these networks may offer better task performance compared to simply using one agent to complete the entire task, they introduce systemic vulnerabilities, such as adversarial composition, semantic misalignment, and cascading operational failures, that existing agent alignment techniques cannot address. In this vision paper, we argue that the trustworthiness of A2A networks cannot be fully guaranteed via retrofitting on existing protocols that are largely designed for individual agents. Rather, it must be architected from the very beginning of the A2A coordination framework. We present a comprehensive conceptual framework that situates trust in A2A systems through four design pillars.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。