测试大模型在多智能体协作中保护隐私的能力,发现主流模型常泄露敏感信息。
MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation
- 构建158个高风险真实场景的评估数据集MAGPIE
- GPT-4o和Claude-2.7-Sonnet误判私密信息比例超25%以上
- 即使有明确指令,仍超50%情况下泄露隐私,适合安全与可信AI研究者
大语言模型驱动的智能体在日程安排、谈判、资源分配等任务中日益依赖多方协作,但其对上下文隐私的理解不足,可能暴露敏感工具与数据库。现有基准多限于单轮简单任务,难以评估复杂情境下的隐私保护能力。本文提出MAGPIE数据集,包含158个跨15个领域的高风险真实场景,要求在不损害任务完成的前提下避免泄露私密信息。实验表明,当前主流模型如GPT-4o和Claude-2.7-Sonnet在识别私密数据时分别错误判定25.2%和43.6%;在多轮对话中,即使收到隐私指令,仍分别在59.9%和50.5%的案例中泄露信息。多智能体系统在71%的场景中无法完成任务。结果揭示当前模型在隐私保护与协作求解之间缺乏对齐。
原文摘要 · Abstract (English)
The proliferation of LLM-based agents has led to increasing deployment of inter-agent collaboration for tasks like scheduling, negotiation, resource allocation etc. In such systems, privacy is critical, as agents often access proprietary tools and domain-specific databases requiring strict confidentiality. This paper examines whether LLM-based agents demonstrate an understanding of contextual privacy. And, if instructed, do these systems preserve inference time user privacy in non-adversarial multi-turn conversation. Existing benchmarks to evaluate contextual privacy in LLM-agents primarily assess single-turn, low-complexity tasks where private information can be easily excluded. We first present a benchmark - MAGPIE comprising 158 real-life high-stakes scenarios across 15 domains. These scenarios are designed such that complete exclusion of private data impedes task completion yet unrestricted information sharing could lead to substantial losses. We then evaluate the current state-of-the-art LLMs on (a) their understanding of contextually private data and (b) their ability to collaborate without violating user privacy. Empirical experiments demonstrate that current models, including GPT-4o and Claude-2.7-Sonnet, lack robust understanding of contextual privacy, misclassifying private data as shareable 25.2\% and 43.6\% of the time. In multi-turn conversations, these models disclose private information in 59.9\% and 50.5\% of cases even under explicit privacy instructions. Furthermore, multi-agent systems fail to complete tasks in 71\% of scenarios. These results underscore that current models are not aligned towards both contextual privacy preservation and collaborative task-solving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。