检测大模型中敏感信息的成员身份风险,发现现有方法在细粒度上效果差。
EL-MIA: Quantifying Membership Inference Risks of Sensitive Entities in LLMs
- 提出EL-MIA框架,聚焦敏感实体的成员身份推断风险。
- 实验显示现有方法在细粒度敏感属性推断上表现不佳。
- 适合关注大模型隐私安全的研究者和开发者参考。
成员身份推断攻击(MIA)旨在判断特定数据点是否属于模型的训练数据。本文提出大模型隐私领域的新任务:针对敏感信息(如个人信息、信用卡号等)的实体级别成员身份风险发现。现有MIA方法可检测完整提示或文档是否在训练集中,但无法捕捉更细粒度的风险。为此,我们提出EL-MIA框架,用于审计大模型中的实体级成员身份风险,并构建了评估该任务的基准数据集。利用该数据集,我们系统比较了现有MIA技术及两种新提出的方法。通过全面分析结果,探讨了实体级别MIA脆弱性与模型规模、训练轮次等表面因素的关系。研究发现,现有MIA方法在敏感属性的细粒度推断上存在局限,而这种脆弱性可通过相对简单的方法识别,凸显出需更强对手来压力测试现有威胁模型的必要性。
原文摘要 · Abstract (English)
Membership inference attacks (MIA) aim to infer whether a particular data point is part of the training dataset of a model. In this paper, we propose a new task in the context of LLM privacy: entity-level discovery of membership risk focused on sensitive information (PII, credit card numbers, etc). Existing methods for MIA can detect the presence of entire prompts or documents in the LLM training data, but they fail to capture risks at a finer granularity. We propose the ``EL-MIA'' framework for auditing entity-level membership risks in LLMs. We construct a benchmark dataset for the evaluation of MIA methods on this task. Using this benchmark, we conduct a systematic comparison of existing MIA techniques as well as two newly proposed methods. We provide a comprehensive analysis of the results, trying to explain the relation of the entity level MIA susceptability with the model scale, training epochs, and other surface level factors. Our findings reveal that existing MIA methods are limited when it comes to entity-level membership inference of the sensitive attributes, while this susceptibility can be outlined with relatively straightforward methods, highlighting the need for stronger adversaries to stress test the provided threat model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。