arXiv:2509.14284cs.CRcs.AI2025-09被引 18

多智能体协作中,普通对话组合可能泄露敏感信息,需新防御机制。

The Sum Leaks More Than Its Parts: Compositional Privacy Risks and Mitigations in Multi-Agent Collaboration

  • 通过建模交互与辅助知识协同放大隐私风险,揭示组合泄露机制。
  • 理论推演防御可阻断97%敏感查询,共识防御平衡性最佳(79.8%)。
  • 适合关注多智能体系统隐私安全的研究者与开发者参考。

随着大语言模型在多智能体系统中的广泛应用,新的隐私风险逐渐浮现,超越了记忆、直接推理或单轮评估的范畴。特别是看似无害的回复,若在多轮交互中累积组合,可能被对手用来恢复敏感信息,我们称之为组合式隐私泄露。本文首次系统研究此类泄露及其缓解方法。首先构建框架,揭示辅助知识与智能体交互如何共同放大隐私风险,即使单次响应本身无害。其次提出两种防御策略:(1) 理论思维防御(ToM),让防御方智能体预测提问者意图,预判输出被滥用的可能性;(2) 协作共识防御(CoDef),响应方智能体基于共享聚合状态投票,限制敏感信息传播。实验在暴露敏感信息与良性推理之间进行平衡评估。结果表明,仅链式思考防护有限(敏感阻断率约39%),而ToM显著提升敏感查询阻断(最高达97%),但会降低良性任务成功率;CoDef取得最佳平衡,实现79.8%的综合表现,凸显显式推理与协作防御结合的优势。本研究揭示了协作式大模型部署中的新型风险,并提供可操作的防护设计思路。

原文摘要 · Abstract (English)

As large language models (LLMs) become integral to multi-agent systems, new privacy risks emerge that extend beyond memorization, direct inference, or single-turn evaluations. In particular, seemingly innocuous responses, when composed across interactions, can cumulatively enable adversaries to recover sensitive information, a phenomenon we term compositional privacy leakage. We present the first systematic study of such compositional privacy leaks and possible mitigation methods in multi-agent LLM systems. First, we develop a framework that models how auxiliary knowledge and agent interactions jointly amplify privacy risks, even when each response is benign in isolation. Next, to mitigate this, we propose and evaluate two defense strategies: (1) Theory-of-Mind defense (ToM), where defender agents infer a questioner's intent by anticipating how their outputs may be exploited by adversaries, and (2) Collaborative Consensus Defense (CoDef), where responder agents collaborate with peers who vote based on a shared aggregated state to restrict sensitive information spread. Crucially, we balance our evaluation across compositions that expose sensitive information and compositions that yield benign inferences. Our experiments quantify how these defense strategies differ in balancing the privacy-utility trade-off. We find that while chain-of-thought alone offers limited protection to leakage (~39% sensitive blocking rate), our ToM defense substantially improves sensitive query blocking (up to 97%) but can reduce benign task success. CoDef achieves the best balance, yielding the highest Balanced Outcome (79.8%), highlighting the benefit of combining explicit reasoning with defender collaboration. Together, our results expose a new class of risks in collaborative LLM deployments and provide actionable insights for designing safeguards against compositional, context-driven privacy leakage.

隐私安全多智能体大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。