针对异构智能体网络,提出一种能伪装成正常更新的图结构毒化攻击。
Graph Representation-based Model Poisoning on the Heterogeneous Internet of Agents
- 基于良性更新构建图结构,用变分图自编码器捕捉依赖关系
- 生成的恶意更新在统计特性上与正常更新一致,准确率显著下降
- 可绕过现有防御机制,适合研究联邦学习安全的学者参考
智能体互联网(IoA)构想一个统一的、以智能体为中心的范式,使异构大语言模型(LLM)智能体能够大规模互联协作。在此范式下,联邦微调(FFT)是关键使能技术,允许分布式LLM智能体在不集中本地数据集的情况下共同训练一个全局智能体。然而,依赖FFT的IoA系统仍易受模型毒化攻击,攻击者可通过上传恶意更新破坏聚合后全局模型的性能。本文提出一种基于图表示的模型毒化(GRMP)攻击,利用监听到的良性更新构建特征相关性图,并采用变分图自编码器捕捉结构依赖关系,生成恶意更新。设计了一种基于增广拉格朗日和次梯度下降的新型攻击算法,优化恶意更新使其保持良性统计特性的同时嵌入对抗目标。实验结果表明,所提GRMP攻击能在不同LLM模型上显著降低准确率,同时保持与良性更新的统计一致性,从而逃避现有防御机制检测,凸显对前瞻性IoA范式的严重威胁。
原文摘要 · Abstract (English)
Internet of Agents (IoA) envisions a unified, agent-centric paradigm where heterogeneous large language model (LLM) agents can interconnect and collaborate at scale. Within this paradigm, federated fine-tuning (FFT) serves as a key enabler that allows distributed LLM agents to co-train an intelligent global LLM without centralizing local datasets. However, the FFT-enabled IoA systems remain vulnerable to model poisoning attacks, where adversaries can upload malicious updates to the server to degrade the performance of the aggregated global LLM. This paper proposes a graph representation-based model poisoning (GRMP) attack, which exploits overheard benign updates to construct a feature correlation graph and employs a variational graph autoencoder to capture structural dependencies and generate malicious updates. A novel attack algorithm is developed based on augmented Lagrangian and subgradient descent methods to optimize malicious updates that preserve benign-like statistics while embedding adversarial objectives. Experimental results show that the proposed GRMP attack can substantially decrease accuracy across different LLM models while remaining statistically consistent with benign updates, thereby evading detection by existing defense mechanisms and underscoring a severe threat to the ambitious IoA paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。