用大模型生成隐私数据,能保留统计特征吗?
Evaluating LLM Simulators as Differentially Private Data Generators
- 用差分隐私数据训练大模型模拟用户行为
- 在ε=1时检测欺诈的AUC达0.70
- 但模型会因自身偏见扭曲时间与人口统计分布
基于大模型的模拟器为生成复杂合成数据提供了新路径,尤其适用于传统差分隐私方法在高维用户画像上表现不佳的场景。本文通过PersonaLedger——一个以差分隐私保护的合成用户画像为种子的金融领域代理模拟器,评估了其对真实统计分布的还原能力。实验发现,在ε=1时,该模拟器在欺诈检测任务中达到0.70的AUC,表现出良好实用性;但存在显著分布漂移问题,源于大模型固有偏见(如学习到的先验)覆盖了输入统计数据,对时间序列和人口统计特征产生系统性偏差。这一缺陷表明,若不解决模型偏见问题,大模型难以胜任需要丰富用户表征的复杂数据生成任务。
原文摘要 · Abstract (English)
LLM-based simulators offer a promising path for generating complex synthetic data where traditional differentially private (DP) methods struggle with high-dimensional user profiles. But can LLMs faithfully reproduce statistical distributions from DP-protected inputs? We evaluate this using PersonaLedger, an agentic financial simulator, seeded with DP synthetic personas derived from real user statistics. We find that PersonaLedger achieves promising fraud detection utility (AUC 0.70 at epsilon=1) but exhibits significant distribution drift due to systematic LLM biases--learned priors overriding input statistics for temporal and demographic features. These failure modes must be addressed before LLM-based methods can handle the richer user representations where they might otherwise excel.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。