arXiv:2506.10171cs.CRcs.AI2025-06被引 6

提出新框架检测大模型代理的隐性隐私泄露风险

Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents

  • 设计多轮迭代探测策略,模拟真实对话中的隐私渗透
  • 发现现有单轮防御无法阻止渐进式信息泄露
  • 提供可量化的审计流程与公开基准测试集

大型语言模型代理正被用于个人助理、客服机器人和临床辅助等场景,虽带来显著效率提升,但持续访问敏感数据也增加了未经授权信息泄露的风险。此类泄露不仅限于显式披露,还可能通过渐进操控或侧信道方式发生。本文提出一种名为「对话隐私泄露操控」(CMPL)的审计框架,通过迭代探测策略压力测试严格遵守隐私指令的代理系统。不同于仅关注单次泄露或显式暴露,CMPL模拟真实的多轮对话,系统性揭示潜在漏洞。在多个领域、数据模态及安全配置下的评估表明,该框架能发现现有单轮防御无法阻断的隐私风险,并深入分析泄露的时间动态、自适应攻击者策略及其对敏感目标的认知演变。研究还提供了基于可量化风险指标的审计流程和开放的对话隐私评测基准。

原文摘要 · Abstract (English)

LLM agents have begun to appear as personal assistants, customer service bots, and clinical aides. While these applications deliver substantial operational benefits, they also require continuous access to sensitive data, which increases the likelihood of unauthorized disclosures. Moreover, these disclosures go beyond mere explicit disclosure, leaving open avenues for gradual manipulation or sidechannel information leakage. This study proposes an auditing framework for conversational privacy that quantifies an agent's susceptibility to these risks. The proposed Conversational Manipulation for Privacy Leakage (CMPL) framework is designed to stress-test agents that enforce strict privacy directives against an iterative probing strategy. Rather than focusing solely on a single disclosure event or purely explicit leakage, CMPL simulates realistic multi-turn interactions to systematically uncover latent vulnerabilities. Our evaluation on diverse domains, data modalities, and safety configurations demonstrates the auditing framework's ability to reveal privacy risks that are not deterred by existing single-turn defenses, along with an in-depth longitudinal study of the temporal dynamics of leakage, strategies adopted by adaptive adversaries, and the evolution of adversarial beliefs about sensitive targets. In addition to introducing CMPL as a diagnostic tool, the paper delivers (1) an auditing procedure grounded in quantifiable risk metrics and (2) an open benchmark for evaluation of conversational privacy across agent implementations.

隐私安全大模型审计对抗测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。