arXiv:2511.17671cs.CRcs.AI2025-11被引 1

多用户语言代理易受跨用户污染攻击,导致无辜用户被误导执行恶意操作。

MURMUR: Using cross-user chatter to break collaborative language agents in groups

  • 用大模型生成真实对话,模拟多人协作场景下的交互行为。
  • 攻击成功率高且影响持续,能触发多个任务中的恶意行为。
  • 提出基于任务聚类的防御方案,适用于共享工作空间的代理系统。

语言代理正从单用户助手转向共享工作区中的多用户协作模式。然而,当前语言模型缺乏隔离用户交互与并发任务的机制,由此产生一种新型攻击向量:跨用户污染(CUP)。在CUP攻击中,攻击者注入看似正常的消息,污染持久的共享状态,进而诱使代理在后续任务中为良性用户执行攻击者指定的非法操作。我们已在真实系统中验证该攻击,成功攻破主流多用户代理。为此,我们提出MURMUR框架,利用大模型生成具有历史感知的真实用户交互,将单用户任务组合成并发的群组场景,系统性研究该现象。实验发现CUP攻击成功率极高,且影响可跨多个任务持续存在,对多用户语言模型部署构成根本性威胁。最后,我们提出基于任务聚类的初步防御策略,以缓解此类新漏洞。

原文摘要 · Abstract (English)

Language agents are rapidly expanding from single-user assistants to multi-user collaborators in shared workspaces and groups. However, today's language models lack a mechanism for isolating user interactions and concurrent tasks, creating a new attack vector inherent to this new setting: cross-user poisoning (CUP). In a CUP attack, an adversary injects ordinary-looking messages that poison the persistent, shared state, which later triggers the agent to execute unintended, attacker-specified actions on behalf of benign users. We validate CUP on real systems, successfully attacking popular multi-user agents. To study the phenomenon systematically, we present MURMUR, a framework that composes single-user tasks into concurrent, group-based scenarios using an LLM to generate realistic, history-aware user interactions. We observe that CUP attacks succeed at high rates and their effects persist across multiple tasks, thus posing fundamental risks to multi-user LLM deployments. Finally, we introduce a first-step defense with task-based clustering to mitigate this new class of vulnerability

语言代理安全攻击多用户协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。