arXiv:2411.07009cs.LGcs.DB2024-11被引 4

提出HCTGAN模型,高效生成具有复杂关系的多表合成数据。

Hierarchical Conditional Tabular GAN for Multi-Tabular Synthetic Data Generation

  • 采用分层条件生成架构,建模多表间复杂依赖关系。
  • 在深度多表数据上生成速度比HMA1快,且始终保证引用完整性。
  • 适合大规模多表数据生成,小数据集可选HMA1保质量。

合成数据生成是解决真实数据受限或隐私合规问题的有效手段。尽管单表数据合成研究已较成熟,但针对具有复杂表间关系的多表数据的研究仍有限。本文提出HCTGAN算法,用于从复杂多表数据中生成合成数据,并与概率模型HMA1进行对比。结果表明,该算法能更高效地生成大规模合成数据,同时保持良好数据质量并始终保证引用完整性。结论指出,HCTGAN适用于生成具有复杂关系的深层多表数据;而当数据量较小时,若侧重数据质量,建议使用HMA1模型。

原文摘要 · Abstract (English)

The generation of synthetic data is a state-of-the-art approach to leverage when access to real data is limited or privacy regulations limit the usability of sensitive data. A fair amount of research has been conducted on synthetic data generation for single-tabular datasets, but only a limited amount of research has been conducted on multi-tabular datasets with complex table relationships. In this paper we propose the algorithm HCTGAN to synthesize multi-tabular data from complex multi-tabular datasets. We compare our results to the probabilistic model HMA1. Our findings show that our proposed algorithm can more efficiently sample large amounts of synthetic data for deep and complex multi-tabular datasets, whilst achieving adequate data quality and always guaranteeing referential integrity. We conclude that the HCTGAN algorithm is suitable for generating large amounts of synthetic data efficiently for deep multi-tabular datasets with complex relationships. We additionally suggest that the HMA1 model should be used on smaller datasets when emphasis is on data quality.

多表生成合成数据图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。