系统评估大模型隐私漏洞,发现攻击效果因设计而异。
On the Privacy of LLMs: An Ablation Study

- 构建统一威胁模型,对比多种隐私攻击在不同配置下的表现。
- 掩码型成员推理攻击信号强且稳定,后门攻击成功率高。
- 针对个人敏感信息的攻击虽准确率低,但风险仍不容忽视。
大型语言模型(LLMs)越来越多地应用于交互式与检索增强场景,引发严重隐私担忧。尽管成员推理(MIA)、属性推理(AIA)、数据提取(DEA)和后门攻击(BA)等已受关注,但它们通常被孤立研究,缺乏对常见系统因素影响的理解。本文提出统一威胁模型与符号体系,复现代表性隐私攻击,并开展系统的消融实验,评估模型架构、规模、数据集特征及检索配置等因素的影响。分析显示:掩码型成员推理攻击表现出强烈且可靠的信号;后门攻击因触发机制实现一致高成功率;而属性推理与数据提取攻击仍具挑战性,准确率较低,但因其针对敏感个人信息,仍构成重大风险。结果表明,大模型系统的隐私风险高度依赖上下文,由设计选择驱动,强调需进行整体评估与审慎部署。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in interactive and retrieval-augmented settings, raising significant privacy concerns. While attacks such as Membership Inference (MIA), Attribute Inference (AIA), Data Extraction (DEA), and Backdoor Attacks (BA) have been studied, they are typically analyzed in isolation, leaving a gap in understanding their behavior under common system factors. In this paper, we introduce a unified threat model and notation, reproduce a representative set of privacy attacks, and conduct a structured ablation study to evaluate the impact of key factors such as model architecture, scale, dataset characteristics, and retrieval configuration. Our analysis reveals clear differences across attack types. Membership inference attacks, particularly mask-based variants, exhibit strong and reliable signals, while backdoor attacks achieve consistently high success rates due to their trigger-based nature. In contrast, attribute inference and data extraction attacks remain more challenging, resulting in lower accuracy, yet they pose significant risks as they target sensitive personal information. Overall, these results highlight that privacy risks in LLM systems are highly context-dependent and driven by design choices, emphasizing the need for holistic evaluation and informed deployment practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。