arXiv:2409.17684cs.LGcs.AI2024-09被引 5

现有合成表格数据生成方法难以保持属性间依赖关系。

Preserving logical and functional dependencies in synthetic tabular data

  • 提出逻辑依赖新概念并设计量化方法
  • 发现主流生成模型无法完整保留函数依赖
  • 适合关注数据隐私与合成质量的研究者

表格数据中属性间的依赖关系普遍存在。然而,现有表格数据生成算法在生成合成数据时是否保留这些依赖关系仍不明确。本文除了已有函数依赖外,引入了属性间的逻辑依赖新概念,并提供量化方法。利用该度量,我们对比了多个最先进的合成数据生成算法在多个公开数据集上的表现。结果表明,当前合成表格数据生成算法在生成过程中未能完全保留函数依赖;但部分模型能够较好保持属性间的逻辑依赖。研究揭示了开发面向特定任务的合成表格数据生成模型的迫切需求与机遇。

原文摘要 · Abstract (English)

Dependencies among attributes are a common aspect of tabular data. However, whether existing tabular data generation algorithms preserve these dependencies while generating synthetic data is yet to be explored. In addition to the existing notion of functional dependencies, we introduce the notion of logical dependencies among the attributes in this article. Moreover, we provide a measure to quantify logical dependencies among attributes in tabular data. Utilizing this measure, we compare several state-of-the-art synthetic data generation algorithms and test their capability to preserve logical and functional dependencies on several publicly available datasets. We demonstrate that currently available synthetic tabular data generation algorithms do not fully preserve functional dependencies when they generate synthetic datasets. In addition, we also showed that some tabular synthetic data generation models can preserve inter-attribute logical dependencies. Our review and comparison of the state-of-the-art reveal research needs and opportunities to develop task-specific synthetic tabular data generation models.

合成数据表格生成数据依赖隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。