测试大模型在多人对话中泄露隐私的风险,发现漏洞比单人场景严重得多。
MuPPET: A Benchmark for Contextual Privacy of LLM Assistants in Multi-Party Conversations

- 构建多角色对话隐私测试基准MuPPET,评估模型在群聊中的隐私暴露风险。
- 实验显示模型在群聊中泄露信息量远超一对一场景,顶尖模型也难幸免。
- 小模型本地部署虽看似安全,但隐私风险更高,现有防护手段效果有限。
大语言模型代理正越来越多地部署于多人环境,代表用户处理敏感个人信息,例如在群聊中。当此类代理泄露私密信息时,所有群成员会同时收到。这一风险在结构上比一对一场景更难控制,因为每条私密信息都需对所有接收者合适。然而,现有上下文隐私评测基准均仅考虑单对话者场景,未能衡量多人隐私风险。我们提出MuPPET(多角色隐私暴露测试),用于评估多角色对话中的上下文隐私。实验表明,模型在多人设置下的信息泄露程度显著高于一对一评估所揭示的水平。前沿模型存在漏洞,而常被用于本地部署以保护敏感数据的小型开源权重模型风险更高。现有上下文隐私防御措施仅提供部分保护,降低模型实用性,且无法解决根本的发言者追踪问题。
原文摘要 · Abstract (English)
LLM agents are increasingly deployed in multi-party environments, handling sensitive personal data on behalf of individual users, for instance in group chats. When such an agent discloses private information, it reaches every group member at once. This risk is structurally harder to control than in one-to-one settings, as every piece of private information must be appropriate for every recipient in the group. Yet all existing contextual privacy benchmarks consider only single-interlocutor settings, leaving multi-party privacy risks unmeasured. We introduce MuPPET (Multi-Party Privacy Exposure Testing), a benchmark for contextual privacy in multi-party conversations. Our experiments show that models leak substantially more in multi-party settings than one-to-one evaluations suggest. Frontier models are vulnerable, and smaller open-weights models, often preferred for local deployment with sensitive data, even more so. Existing contextual privacy defences offer only partial protection, degrade utility, and do not resolve the underlying party-tracking problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。