arXiv:2409.01380cs.CRcs.CL2024-09被引 68

首次针对上下文学习设计会员推断攻击,仅用生成文本即可高精度识别数据是否在训练集中。

Membership Inference Attacks Against In-Context Learning

  • 仅凭生成文本,不依赖概率信息,设计四类适配不同场景的攻击策略。
  • 对LLaMA等模型攻击准确率超95%,显著高于传统基于概率的攻击。
  • 提出混合攻击与多维防御方案,为大模型隐私安全提供新思路。

将大型语言模型(LLMs)适配特定任务会带来计算效率问题,促使人们探索如上下文学习(ICL)等高效方法。然而,在现实假设下,ICL对隐私攻击的脆弱性尚未得到充分研究。本文首次提出专为ICL设计的会员推断攻击,仅依赖生成文本而无需其关联概率。我们提出了四种适配不同受限场景的攻击策略,并在四个主流大语言模型上开展广泛实验。结果表明,我们的攻击在大多数情况下能准确判断成员身份,例如对LLaMA的准确率优势达95%,说明其风险远高于现有基于概率的攻击。此外,我们提出一种融合前述策略的混合攻击,在多数情况下实现超过95%的准确率优势。同时,我们研究了针对数据、指令和输出三方面的潜在防御措施,结果表明,结合来自不同维度的防御可显著降低隐私泄露,提升隐私保障水平。

原文摘要 · Abstract (English)

Adapting Large Language Models (LLMs) to specific tasks introduces concerns about computational efficiency, prompting an exploration of efficient methods such as In-Context Learning (ICL). However, the vulnerability of ICL to privacy attacks under realistic assumptions remains largely unexplored. In this work, we present the first membership inference attack tailored for ICL, relying solely on generated texts without their associated probabilities. We propose four attack strategies tailored to various constrained scenarios and conduct extensive experiments on four popular large language models. Empirical results show that our attacks can accurately determine membership status in most cases, e.g., 95\% accuracy advantage against LLaMA, indicating that the associated risks are much higher than those shown by existing probability-based attacks. Additionally, we propose a hybrid attack that synthesizes the strengths of the aforementioned strategies, achieving an accuracy advantage of over 95\% in most cases. Furthermore, we investigate three potential defenses targeting data, instruction, and output. Results demonstrate combining defenses from orthogonal dimensions significantly reduces privacy leakage and offers enhanced privacy assurances.

隐私安全大模型会员推断上下文学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。