arXiv:2606.07150cs.CRcs.AI2026-06

揭露智能体通信图谱泄露风险,提出保护工作流完整性的新防御机制

From Privacy to Workflow Integrity: Communication-Graph Metadata in Autonomous Agent Interoperability

论文配图:From Privacy to Workflow Integrity: Communication-Graph Metadata in Autonomous Agent Interoperability
图 1 · 摘自论文原文
  • 构建通信图谱威胁模型,揭示元数据暴露对工作流的致命影响
  • 被动分析仅靠初始交互即可6倍于随机水平识别任务类型
  • 防御策略需兼顾隐私与工作流完整性,二者在对抗中可分离

A2A和MCP等智能体互操作协议虽标准化了智能体间的通信内容,但依赖基于地址的传输方式。无论通过HTTP(S)还是基于MLS的SLIM等保护性绑定,这些传输均能加密消息内容,却使通信图谱暴露:谁在何时与谁通信、频率如何。在智能体系统中,该图谱比隐私泄露更严重——端点具能力标签,工作流结构化且链式执行,交互与动作强耦合,攻击者可从初始行为快速识别待执行的工作流,并在机器速度下抢先干预。威胁本质是工作流完整性受损,而非仅隐私问题。本文提出通信图谱威胁模型,指出其关键危害在于跨信任域暴露与自主行动的结合。定义传输层与启动层隐私属性,采用不可区分性博弈语义进行形式化,评估多种传输方式,并通过真实多智能体A2A流量案例研究发现,一个无标签分类器仅凭被动元数据就能在6倍于随机水平下识别任务类别,甚至仅凭首次交互;即使防御意识强的对手也无法完全消除此风险,唯有完整属性组合才能将识别率拉回随机水平。攻击获利不等于可恢复性:在固定预算下,对手获得0.63倍全知攻击者的收益(其中0.41来自工作流初始阶段),且受最高精度项主导,表明完整性与隐私在防御条件下可分离。

原文摘要 · Abstract (English)

Agent-interoperability protocols such as A2A and MCP standardize what agents say to one another but assume address-based transport. Whether over HTTP(S) or a content-protecting binding such as MLS-based SLIM, these transports protect message content yet leave the communication graph exposed: which agent contacts which, when, and how often. In agent systems this graph is more consequential than a privacy framing suggests. Endpoints are capability-labeled, workflows are structured and chained, and interactions are coupled to actions, so an observer recovers more than past relationships: it can recognize a recurring pending workflow from its opening and, at machine speed, act on it before it completes. The threat is one of workflow integrity, not privacy alone. We give a threat model for the communication graph and locate what makes its metadata distinctively consequential: not stronger fingerprinting but exposure across independent trust domains, coupled to autonomous action. We define transport- and bootstrap-layer privacy properties, give them an indistinguishability-game semantics, evaluate transports, and give an A2A case study where a metadata-protecting binding surfaces its implicit identity assumptions. On a corpus of real multi-agent A2A traffic from the official reference agents, on a live A2A binding, and with a generative model as a controlled instrument, a label-blind classifier recovers a task's class from passive metadata at 6x chance, and from only its opening; a defense-aware adversary does not overturn this, and only the full set of properties drives recovery toward chance. Acting on the leak is distinct from recoverability: under a fixed budget an adversary captures 0.63 of a clairvoyant attacker's advantage on the corpus (0.41 from a workflow's opening), governed by top-ranked precision rather than overall accuracy, so integrity and privacy come apart under defense.

智能体系统工作流安全通信元数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。