AI agents互交催生新安全威胁,这篇论文系统定义了多智能体安全领域。
Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents
- 提出多智能体安全新领域,聚焦智能体交互产生的新型威胁
- 梳理出由协同攻击、隐私泄露等引发的系统性风险图谱
- 为安全研究提供统一框架,适合关注AI治理与系统安全者阅读
AI智能体正直接在互联网平台和物理环境中相互交互,带来超越传统网络安全与人工智能安全框架的新挑战。自由形式的协议虽促进任务泛化,却也引发秘密串通与协同群体攻击等新威胁。网络效应会快速传播隐私泄露、虚假信息、越狱攻击和数据污染,而多智能体的分散部署与隐匿优化使攻击者更难被监管,形成系统级持久威胁。尽管这些挑战至关重要,相关研究仍分散于人工智能安全、多智能体学习、复杂系统、网络安全、博弈论、分布式系统和技术治理等多个领域。本文提出多智能体安全这一新领域,致力于防范智能体间交互(包括直接或通过共享环境)所引发或放大的威胁,涵盖对其他智能体、人类及制度的影响,并揭示分布式与去中心化场景下的安全-效用、安全-安全权衡。初步工作包括:(1) 对交互智能体引发的威胁图景进行分类;(2) 将跨子领域的研究成果应用于多智能体安全;(3) 提出统一研究议程以应对设计安全智能体系统与交互环境的关键挑战。通过识别这些缺口,旨在引导该关键领域研究,释放大规模智能体部署的社会经济潜力,增强公众信任,并缓解关键基础设施与国防领域的国家安全风险。
原文摘要 · Abstract (English)
AI agents are beginning to interact with each other directly and across internet platforms and physical environments, creating security challenges beyond traditional cybersecurity and AI safety frameworks. Free-form protocols are essential for AI's task generalization but enable new threats like secret collusion and coordinated swarm attacks. Network effects can rapidly spread privacy breaches, disinformation, jailbreaks, and data poisoning, while multi-agent dispersion and stealth optimization help adversaries evade oversight - creating novel persistent threats at a systemic level. Despite their critical importance, these security challenges remain understudied, with research fragmented across disparate fields including AI security, multi-agent learning, complex systems, cybersecurity, game theory, distributed systems, and technical AI governance. We introduce multi-agent security, a new field dedicated to securing networks of AI agents against threats that emerge or amplify through their interactions - whether direct or indirect via shared environments - with each other, humans, and institutions, and characterise fundamental security-utility and security-security trade-offs across both distributed and decentralised settings. Our preliminary work (1) taxonomizes the threat landscape arising from interacting AI agents, (2) offers applications to multi-agent security for work across diffuse subfields, and (3) proposes a unified research agenda addressing open challenges in designing secure agent systems and interaction environments. By identifying these gaps, we aim to guide research in this critical area to unlock the socioeconomic potential of large-scale agent deployment, foster public trust, and mitigate national security risks in critical infrastructure and defense contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。