研究监控下多智能体对话的隐蔽攻击,发现仅靠监控无法杜绝风险。
Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discussions
- 设计新攻击方法应对持续监控环境
- 实验表明监控下仍可成功实施隐蔽攻击
- 警示安全系统需超越单纯监控
多智能体讨论被广泛应用,促使攻击方法研究不断深入。本文研究一种实际但未被充分探索的攻击场景——讨论监控场景,即异常检测器持续监控智能体间通信并拦截恶意消息。尽管现有攻击在无监控时有效,但其行为模式可被检测,在监控下基本失效。然而,这是否意味着监控足以保障安全?我们提出一种专为监控场景设计的新攻击方法。大量实验表明,即便在持续监控下,有效攻击依然可行,说明仅靠监控无法消除对抗风险。
原文摘要 · Abstract (English)
Multi-agent discussions have been widely adopted, motivating growing efforts to develop attacks that expose their vulnerabilities. In this work, we study a practical yet largely unexplored attack scenario, the discussion-monitored scenario, where anomaly detectors continuously monitor inter-agent communications and block detected adversarial messages. Although existing attacks are effective without discussion monitoring, we show that they exhibit detectable patterns and largely fail under such monitoring constraints. But does this imply that monitoring alone is sufficient to secure multi-agent discussions? To answer this question, we develop a novel attack method explicitly tailored to the discussion-monitored scenario. Extensive experiments demonstrate that effective attacks remain possible even under continuous monitoring, indicating that monitoring alone does not eliminate adversarial risks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。