发现大模型多智能体系统中的隐蔽作恶行为,并提出心理检测框架及时识别。
Who's the Mole? Modeling and Detecting Intention-Hiding Malicious Agents in LLM-Based Multi-Agent Systems
- 设计四种隐蔽攻击模式,干扰任务但难被发现。
- 在六大数据集上验证,攻击成功率高且能绕过现有防御。
- 基于人格模型和审讯技术,主动识别潜在恶意智能体。
由大语言模型驱动的多智能体系统(LLM-MAS)在协作解决问题方面展现出强大能力,但其部署也带来了新的安全风险。现有研究主要关注单智能体场景,而对多智能体系统的安全性仍缺乏探索。为此,本文系统研究了LLM-MAS中的意图隐藏威胁,设计了四种代表性攻击范式,在集中式、去中心化和分层通信结构下评估其效果。实验表明,这些攻击极具破坏性且易于规避现有防御机制。为应对该问题,我们提出AgentXposed——一种受心理学启发的检测框架,融合HEXACO人格模型与瑞德审讯法,通过渐进式问卷探测与行为监控相结合,实现对恶意智能体的主动识别。在六大数据集上,针对本文攻击及两种基线威胁的实验表明,AgentXposed能有效检测多种恶意行为,在多种通信设置下均表现稳健。
原文摘要 · Abstract (English)
Multi-agent systems powered by Large Language Models (LLM-MAS) have demonstrated remarkable capabilities in collaborative problem-solving. However, their deployment also introduces new security risks. Existing research on LLM-based agents has primarily examined single-agent scenarios, while the security of multi-agent systems remains largely unexplored. To address this gap, we present a systematic study of intention-hiding threats in LLM-MAS. We design four representative attack paradigms that subtly disrupt task completion while maintaining a high degree of stealth, and evaluate them under centralized, decentralized, and layered communication structures. Experimental results show that these attacks are highly disruptive and can easily evade existing defense mechanisms. To counter these threats, we propose AgentXposed, a psychology-inspired detection framework. AgentXposed draws on the HEXACO personality model, which characterizes agents through psychological trait dimensions, and the Reid interrogation technique, a structured method for eliciting concealed intentions. By combining progressive questionnaire probing with behavior-based inter-agent monitoring, the framework enables the proactive identification of malicious agents before harmful actions are carried out. Extensive experiments across six datasets against both our proposed attacks and two baseline threats demonstrate that AgentXposed effectively detects diverse forms of malicious behavior, achieving strong robustness across multiple communication settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。