arXiv:2605.27766cs.AI2026-05

多智能体社交环境会诱发隐私泄露,即使有防护也难避免。

Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems

论文配图:Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems
图 1 · 摘自论文原文
  • 构建千智能体社交模拟平台,测试长期互动中的隐私安全
  • 多轮交互使隐私泄露率从19.95%升至45.30%,社交传染效应显著
  • 即使有隐私指令,泄露率仍超37.8%,适合关注智能体安全的研究者

LLM安全评估通常在孤立场景下进行,但实际部署的AI智能体常在持续社交环境中与其他智能体共存。本文引入类似Moltbook的仿真平台,让数千个LLM智能体在模拟一个月内跨社群互动,评估不同社会压力下的隐私安全问题。结果发现,从单轮对话转向多轮社交评估后,隐私泄露率在OpenAI模型中从CIMemories的19.95%上升至本文方法的45.30%;信息泄露具有社交传染性,观察到同伴泄露后,智能体泄露概率提升8倍;即使施加显式隐私指令,泄露率仍高于37.8%。研究表明,静态对话基准严重低估了代理部署中的风险,仅凭社交上下文就足以引发单轮评估无法发现的敏感信息泄露。

原文摘要 · Abstract (English)

LLM safety evaluations predominantly test models in isolation, yet deployed AI agents increasingly operate within persistent social environments alongside other agents. We introduce a Moltbook-style simulation platform where thousands of LLM agents interact across communities over a simulated month, and use it to evaluate privacy as a downstream safety concern under varying degrees of social pressure. We find that shifting from single turn to multi turn social evaluation amplifies privacy violations (CIMemories 19.95% to Ours 45.30% across OpenAI models), that leakage is socially contagious, with agents 8 times more likely to disclose sensitive information after observing a peer do so, and that explicit privacy instructions reduce but do not eliminate this effect, leaving leakage rates above 37.8% even with safeguards. Our findings suggest that static chat based safety benchmarks systematically underestimate risks in agentic deployment, and that social context alone is sufficient to elicit sensitive disclosures that single turn evaluations would never surface.

隐私安全多智能体大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。