用最大熵原理生成表格数据,高效且准确。
GEM-T: Generative Tabular Data via Fitting Moments
- 基于最大熵原理直接建模列间高阶关联。
- 在34个数据集中有23个超越或媲美顶尖深度模型。
- 参数量极低,适合敏感或小规模表格数据生成。
表格数据主导数据科学,但生成模型面临数据有限或敏感的挑战。本文提出一种基于最大熵(MaxEnt)的新方法GEM-T(生成熵最大化表),直接捕捉训练数据中任意阶次的列间交互关系(如成对、三阶等)。在34个公开数据集上广泛测试,GEM-T在23个数据集中表现达到或超过以往被视为最先进的深度神经网络方法(占比68%)。尤为关键的是,GEM-T的可训练参数量减少数个数量级,表明真实世界数据中的大部分信息存在于低维、可能可解释的相关性中,前提是输入数据经过适当变换。此外,最大熵方法更擅长处理异构数据类型(连续、离散、分类)、缺乏局部结构等表格数据特征。GEM-T为结构化数据轻量高效生成模型提供了有前景的方向。
原文摘要 · Abstract (English)
Tabular data dominates data science but poses challenges for generative models, especially when the data is limited or sensitive. We present a novel approach to generating synthetic tabular data based on the principle of maximum entropy -- MaxEnt -- called GEM-T, for ``generative entropy maximization for tables.'' GEM-T directly captures nth-order interactions -- pairwise, third-order, etc. -- among columns of training data. In extensive testing, GEM-T matches or exceeds deep neural network approaches previously regarded as state-of-the-art in 23 of 34 publicly available datasets representing diverse subject domains (68\%). Notably, GEM-T involves orders-of-magnitude fewer trainable parameters, demonstrating that much of the information in real-world data resides in low-dimensional, potentially human-interpretable correlations, provided that the input data is appropriately transformed first. Furthermore, MaxEnt better handles heterogeneous data types (continuous vs. discrete vs. categorical), lack of local structure, and other features of tabular data. GEM-T represents a promising direction for light-weight high-performance generative models for structured data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。