攻击者可从Llama 3中提取密码、邮箱等个人隐私信息
Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models
- 用精心设计的提示词查询模型,触发数据泄露
- 成功提取出密码、邮箱、账号等敏感信息
- 提醒小模型也需重视隐私防护,适合安全研究者参考
大型语言模型(LLMs)已彻底改变自然语言处理,但其对训练数据的记忆能力带来了严重的隐私风险。本文研究了Meta开发的多语言LLM Llama 3.2遭受模型反演攻击的可行性。通过精心设计的查询提示,我们成功从模型中提取出个人身份信息(PII),如密码、电子邮件地址和账户号码。结果表明,即使是较小的LLM也易受隐私攻击,凸显了构建稳健防御机制的紧迫性。我们讨论了差分隐私与数据清洗等缓解策略,并呼吁进一步开展隐私保护机器学习的研究。
原文摘要 · Abstract (English)
Large language models (LLMs) have transformed natural language processing, but their ability to memorize training data poses significant privacy risks. This paper investigates model inversion attacks on the Llama 3.2 model, a multilingual LLM developed by Meta. By querying the model with carefully crafted prompts, we demonstrate the extraction of personally identifiable information (PII) such as passwords, email addresses, and account numbers. Our findings highlight the vulnerability of even smaller LLMs to privacy attacks and underscore the need for robust defenses. We discuss potential mitigation strategies, including differential privacy and data sanitization, and call for further research into privacy-preserving machine learning techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。