提出防御大模型网络中顽固攻击者的安全框架
Don't Trust Stubborn Neighbors: A Security Framework for Agentic Networks
- 用社会学观点形成模型分析代理间信任风险
- 单个顽固代理可引发意见级联并控制集体决策
- 动态调整信任度可有效防御攻击且保持协作效率
基于大语言模型的多智能体系统(LLM-MAS)在网页自动化、行程规划和协同解题等任务中广泛应用,但其交互特性引入了新安全风险:恶意或被攻陷的智能体可通过通信渠道传播虚假信息,操纵集体结果。本文借鉴社会学中的弗里德金-约翰森观点形成模型,构建了一个通用理论框架来研究LLM-MAS。实验验证该模型能准确捕捉不同网络结构与攻防场景下的系统行为。理论与实证均表明,单一高度顽固且具有说服力的代理即可主导系统动态,引发意见级联,重塑集体观点。分析揭示三种增强安全性的机制:增加良性代理数量、提高代理内在顽固性或抗同化能力、降低对潜在对手的信任。由于规模扩展成本高且高顽固性削弱共识能力,我们提出一种自适应信任防御机制,动态调节代理间信任以限制敌方影响,同时维持协作性能。大量实验证实该机制能有效抵御操纵攻击。
原文摘要 · Abstract (English)
Large Language Model (LLM)-based Multi-Agent Systems (MASs) are increasingly deployed for agentic tasks, such as web automation, itinerary planning, and collaborative problem solving. Yet, their interactive nature introduces new security risks: malicious or compromised agents can exploit communication channels to propagate misinformation and manipulate collective outcomes. In this paper, we study how such manipulation can arise and spread by borrowing the Friedkin-Johnsen opinion formation model from social sciences to propose a general theoretical framework to study LLM-MAS. Remarkably, this model closely captures LLM-MAS behavior, as we verify in extensive experiments across different network topologies and attack and defense scenarios. Theoretically and empirically, we find that a single highly stubborn and persuasive agent can take over MAS dynamics, underscoring the systems' high susceptibility to attacks by triggering a persuasion cascade that reshapes collective opinion. Our theoretical analysis reveals three mechanisms to increase system security: a) increasing the number of benign agents, b) increasing the innate stubbornness or peer-resistance of agents, or c) reducing trust in potential adversaries. Because scaling is computationally expensive and high stubbornness degrades the network's ability to reach consensus, we propose a new mechanism to mitigate threats by a trust-adaptive defense that dynamically adjusts inter-agent trust to limit adversarial influence while maintaining cooperative performance. Extensive experiments confirm that this mechanism effectively defends against manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。