SentinelNet用信用机制动态检测多智能体中的恶意行为,防患于未然。
SentinelNet: Safeguarding Multi-Agent Collaboration Through Credit-Based Dynamic Threat Detection
- 基于信用的去中心化检测框架,通过对比学习训练代理判断消息可信度。
- 两次辩论回合内接近100%识别恶意代理,恢复95%系统准确率。
- 适合高可靠性要求的多智能体协作场景,如自动驾驶、金融决策。
大型语言模型驱动的多智能体系统(MAS)面临恶意智能体对可靠性和决策能力的威胁。现有防御方法因反应式设计或集中式架构而效果有限,易产生单点故障。为此,我们提出SentinelNet,首个去中心化的主动威胁检测与缓解框架。该框架为每个智能体配备基于信用的检测器,通过对抗性辩论轨迹的对比学习进行训练,实现消息可信度自主评估,并通过底-k剔除动态排序邻居,抑制恶意通信。为解决攻击数据稀缺问题,系统生成模拟多样化威胁的对抗轨迹,确保训练鲁棒性。在MAS基准测试中,SentinelNet在两次辩论回合内近乎完全检测出恶意代理,将受损基线系统的准确率恢复至95%。其跨领域与攻击模式表现出强泛化能力,开创了保障协作式MAS安全的新范式。
原文摘要 · Abstract (English)
Malicious agents pose significant threats to the reliability and decision-making capabilities of Multi-Agent Systems (MAS) powered by Large Language Models (LLMs). Existing defenses often fall short due to reactive designs or centralized architectures which may introduce single points of failure. To address these challenges, we propose SentinelNet, the first decentralized framework for proactively detecting and mitigating malicious behaviors in multi-agent collaboration. SentinelNet equips each agent with a credit-based detector trained via contrastive learning on augmented adversarial debate trajectories, enabling autonomous evaluation of message credibility and dynamic neighbor ranking via bottom-k elimination to suppress malicious communications. To overcome the scarcity of attack data, it generates adversarial trajectories simulating diverse threats, ensuring robust training. Experiments on MAS benchmarks show SentinelNet achieves near-perfect detection of malicious agents, close to 100% within two debate rounds, and recovers 95% of system accuracy from compromised baselines. By exhibiting strong generalizability across domains and attack patterns, SentinelNet establishes a novel paradigm for safeguarding collaborative MAS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。