arXiv:2503.06444cs.LGcs.AI2025-03AAAI

高维表格数据生成难题,用条件控制提升模型稳定性。

Towards Synthesizing High-Dimensional Tabular Data with Limited Samples

  • 引入扰动真实样本作为辅助输入,增强模型对高维低样本数据的适应能力。
  • 在多数据集上平均准确率超越顶尖模型90%以上。
  • 适合处理样本少、维度高的表格数据生成任务。

基于扩散的表格数据生成模型已取得良好效果。然而,当数据维度升高时,现有模型容易退化,甚至表现不如简单的非扩散基模型。这是因为在高维空间中有限的训练样本难以让生成模型准确捕捉数据分布。为缓解学习信号不足并稳定训练,我们提出 CtrTab:一种条件控制的扩散模型,在训练中注入扰动的真实样本作为辅助输入。该设计在模型对控制信号的敏感性上引入隐式 L2 正则化,提升了高维低数据场景下的鲁棒性与稳定性。在多个数据集上的实验表明,CtrTab 显著优于现有最先进模型,平均准确率提升超过 90%。

原文摘要 · Abstract (English)

Diffusion-based tabular data synthesis models have yielded promising results. However, when the data dimensionality increases, existing models tend to degenerate and may perform even worse than simpler, non-diffusion-based models. This is because limited training samples in high-dimensional space often hinder generative models from capturing the distribution accurately. To mitigate the insufficient learning signals and to stabilize training under such conditions, we propose CtrTab, a condition-controlled diffusion model that injects perturbed ground-truth samples as auxiliary inputs during training. This design introduces an implicit L2 regularization on the model's sensitivity to the control signal, improving robustness and stability in high-dimensional, low-data scenarios. Experimental results across multiple datasets show that CtrTab outperforms state-of-the-art models, with a performance gap in accuracy over 90% on average.

表格生成扩散模型高维数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。