针对大模型生成特性设计新攻击,能发现训练数据的隐含记忆。
Context-Aware Membership Inference Attacks against Pre-trained Large Language Models
- 基于子序列困惑度动态设计统计测试
- 显著优于旧方法,揭示上下文依赖的记忆模式
- 适合研究模型隐私与安全的学者
预训练大语言模型(LLMs)的成员推理攻击旨在判断某条数据是否曾用于模型训练。以往针对分类模型的成员推理攻击在大语言模型上失效,因其忽略了大模型跨词元序列的生成特性。本文提出一种新型攻击方法,将成员推理的统计检验适配于数据点内部子序列的困惑度动态变化。实验表明,该方法显著优于现有方法,揭示了预训练大语言模型中存在上下文依赖的记忆模式。
原文摘要 · Abstract (English)
Membership Inference Attacks (MIAs) on pre-trained Large Language Models (LLMs) aim at determining if a data point was part of the model's training set. Prior MIAs that are built for classification models fail at LLMs, due to ignoring the generative nature of LLMs across token sequences. In this paper, we present a novel attack on pre-trained LLMs that adapts MIA statistical tests to the perplexity dynamics of subsequences within a data point. Our method significantly outperforms prior approaches, revealing context-dependent memorization patterns in pre-trained LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。