用上下文学习突破小数据下生成质量与隐私的矛盾
Breaking the Quality-Privacy Tradeoff in Tabular Data Generation via In-Context Learning

- 将表格生成转化为上下文学习任务,利用预训练结构先验
- 在14个真实数据集上同时提升生成质量和隐私保护
- 适合需要高保真且安全合成数据的研究者
表格数据生成旨在生成高质量数据的同时保护隐私。然而,我们发现现有表格生成模型在小数据场景下存在明显质量-隐私权衡:提高数据质量通常导致对训练样本的记忆增强,从而削弱隐私保护。这一权衡源于小规模训练集使特定数据集生成模型难以区分可泛化的结构与样本特异性模式。为此,我们提出DiffICL,将表格数据生成建模为上下文学习问题。不同于从头拟合每个数据集,DiffICL利用从大量数据集中学习到的预训练结构先验,使其能从有限上下文中推断数据分布,而非记忆单个样本。我们在14个真实世界数据集上评估了DiffICL,结果表明其同时提升了数据质量和隐私保护,并生成了有效的数据增强样本。研究结果表明,通过更优的训练范式可缓解质量-隐私权衡。
原文摘要 · Abstract (English)
Tabular data synthesis aims to generate high-quality data while preserving privacy. However, we find that existing tabular generative models exhibit a clear tradeoff in the small-data regime: improving data quality typically comes at the cost of increased memorization of training samples, thereby weakening privacy protection. This tradeoff arises because small training sets make it difficult for dataset-specific generative models to distinguish generalizable structure from sample-specific patterns. To address this, we propose DiffICL, which formulates tabular data generation as an in-context learning problem. Instead of fitting each dataset from scratch,DiffICL leverages pretrained structural priors learned from a large collection of datasets, enabling it to infer data distributions from limited context rather than memorizing individual samples. We evaluate DiffICL on 14 real-world datasets. Results show that DiffICL improves both data quality and privacy, and generate synthetic data that provides effective data augmentation. Our findings suggest that the quality-privacy tradeoff can be improved through better training paradigms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。