构建首个百万级全合成隐私数据集,解决敏感数据研究匮乏问题。
Privasis: Synthesizing the Largest "Public" Private Dataset from Scratch
- 从零构建140万条带丰富隐私属性的合成文本数据
- 训练的40亿参数模型优于GPT-5等大模型的文本脱敏效果
- 适合隐私保护、数据安全与智能代理研究者使用
涉及敏感隐私数据的研究长期受限于数据稀缺,而现代AI代理如OpenClaw和Gemini Agent已能持续访问高度敏感个人信息。为突破这一瓶颈并应对日益增长的风险,我们提出Privasis(隐私绿洲),首个完全从零构建的百万级合成数据集,包含140万条记录,涵盖医疗史、法律文件、财务记录、日程安排及短信等多种文档类型,共标注5510万项隐私属性(如种族、出生日期、工作单位等)。该数据集在规模、质量和多样性上均远超现有数据集。我们利用Privasis构建文本脱敏的平行语料库,通过分解与定向清洗的管道实现高效去标识化。基于此训练的紧凑型脱敏模型(≤40亿参数)性能超越GPT-5和Qwen-3 235B等先进大模型。我们计划开源数据、模型与代码,推动敏感领域研究发展。
原文摘要 · Abstract (English)
Research involving privacy-sensitive data has always been constrained by data scarcity, standing in sharp contrast to other areas that have benefited from data scaling. This challenge is becoming increasingly urgent as modern AI agents--such as OpenClaw and Gemini Agent--are granted persistent access to highly sensitive personal information. To tackle this longstanding bottleneck and the rising risks, we present Privasis (i.e., privacy oasis), the first million-scale fully synthetic dataset entirely built from scratch--an expansive reservoir of texts with rich and diverse private information--designed to broaden and accelerate research in areas where processing sensitive social data is inevitable. Compared to existing datasets, Privasis, comprising 1.4 million records, offers orders-of-magnitude larger scale with quality, and far greater diversity across various document types, including medical history, legal documents, financial records, calendars, and text messages with a total of 55.1 million annotated attributes such as ethnicity, date of birth, workplace, etc. We leverage Privasis to construct a parallel corpus for text sanitization with our pipeline that decomposes texts and applies targeted sanitization. Our compact sanitization models (<=4B) trained on this dataset outperform state-of-the-art large language models, such as GPT-5 and Qwen-3 235B. We plan to release data, models, and code to accelerate future research on privacy-sensitive domains and agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。