构建合成敏感数据集,用于研究大模型记忆隐私信息的风险。
PANORAMA: A synthetic PII-laced dataset for studying sensitive data memorization in LLMs
- 基于9674个虚拟人物生成38万条含真实场景的敏感信息
- 不同内容类型的记忆风险差异明显,重复率越高记忆越强
- 适合隐私安全研究者和模型审计人员使用
大型语言模型(LLMs)对敏感个人身份信息(PII)的记憶带来了日益增长的隐私风险。现有研究因缺乏全面、真实且合乎伦理的数据集而受限。我们提出PANORAMA——一种大规模合成语料库,包含384,789条样本,源自9,674个基于真实人口统计特征构建的合成用户档案。通过零样本提示与OpenAI o3-mini生成维基文章、社交媒体帖子、论坛讨论、评论、市场列表等多样化内容,嵌入情境合理的PII和敏感信息。通过在Mistral-7B上以1×、5×、10×、25×复制率微调子集数据,测量PII记忆率,发现记忆率随重复次数上升,并在不同内容类型间存在差异,验证了该数据集模拟真实记忆风险的能力。数据与代码公开,可用于隐私风险评估、模型审计及隐私保护模型开发。
原文摘要 · Abstract (English)
The memorization of sensitive and personally identifiable information (PII) by large language models (LLMs) poses growing privacy risks as models scale and are increasingly deployed in real-world applications. Existing efforts to study sensitive and PII data memorization and develop mitigation strategies are hampered by the absence of comprehensive, realistic, and ethically sourced datasets reflecting the diversity of sensitive information found on the web. We introduce PANORAMA - Profile-based Assemblage for Naturalistic Online Representation and Attribute Memorization Analysis, a large-scale synthetic corpus of 384,789 samples derived from 9,674 synthetic profiles designed to closely emulate the distribution, variety, and context of PII and sensitive data as it naturally occurs in online environments. Our data generation pipeline begins with the construction of internally consistent, multi-attribute human profiles using constrained selection to reflect real-world demographics such as education, health attributes, financial status, etc. Using a combination of zero-shot prompting and OpenAI o3-mini, we generate diverse content types - including wiki-style articles, social media posts, forum discussions, online reviews, comments, and marketplace listings - each embedding realistic, contextually appropriate PII and other sensitive information. We validate the utility of PANORAMA by fine-tuning the Mistral-7B model on 1x, 5x, 10x, and 25x data replication rates with a subset of data and measure PII memorization rates - revealing not only consistent increases with repetition but also variation across content types, highlighting PANORAMA's ability to model how memorization risks differ by context. Our dataset and code are publicly available, providing a much-needed resource for privacy risk assessment, model auditing, and the development of privacy-preserving LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。