arXiv:2505.23847cs.CRcs.AI2025-05被引 21

跨域多智能体大模型存在七类新型安全风险,需重新构建信任机制。

Seven Security Challenges in Cross-domain Multi-agent LLM Systems

  • 提出跨域协作中因智能体交互引发的七类新安全挑战
  • 揭示单一智能体看似安全却在跨域通信中泄露隐私或违规
  • 为研究人员提供攻击模拟、评估指标和未来方向指引

大语言模型正演变为跨组织协作的自主智能体,支持灾难响应、供应链优化等任务,可在不共享数据的前提下整合分布式知识。然而,跨域合作打破了现有对齐与管控技术所依赖的统一信任假设。一个孤立时无害的智能体,在接收不可信同伴消息后,可能泄露敏感信息或违反政策,其风险源于多智能体交互中涌现的动态行为,而非传统软件漏洞。本文系统梳理跨域多智能体大模型的安全议题,提出七类新型安全挑战,每类均包含可行攻击方式、安全评估指标及未来研究建议。

原文摘要 · Abstract (English)

Large language models (LLMs) are rapidly evolving into autonomous agents that cooperate across organizational boundaries, enabling joint disaster response, supply-chain optimization, and other tasks that demand decentralized expertise without surrendering data ownership. Yet, cross-domain collaboration shatters the unified trust assumptions behind current alignment and containment techniques. An agent benign in isolation may, when receiving messages from an untrusted peer, leak secrets or violate policy, producing risks driven by emergent multi-agent dynamics rather than classical software bugs. This position paper maps the security agenda for cross-domain multi-agent LLM systems. We introduce seven categories of novel security challenges, for each of which we also present plausible attacks, security evaluation metrics, and future research guidelines.

多智能体安全挑战大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。