TabularARGN高效生成高质量表格数据,支持多种复杂场景。
TabularARGN: A Flexible and Efficient Auto-Regressive Framework for Generating High-Fidelity Synthetic Data
- 基于自回归框架,统一处理混合类型与序列数据
- 在多个基准上达到领先合成数据质量,训练推理更快
- 适合需要公平性、缺失值填补或条件生成的工业应用
表格数据的合成需在保真度、效率和泛化能力间取得平衡。我们提出表格式自回归生成网络(TabularARGN),一个可处理混合类型、多变量及序列数据的灵活框架。通过学习所有可能的条件概率,该框架支持公平性感知生成、缺失值填补以及任意列子集的条件生成。在包含复杂关系的真实数据集上的评估表明,该方法在保持顶尖合成数据质量的同时,显著降低训练与推理时间,适用于结构多样的大规模数据集。通过兼顾灵活性与性能,该框架为各行业实用化合成数据生成提供了新路径。
原文摘要 · Abstract (English)
Synthetic data generation for tabular datasets must balance fidelity, efficiency, and versatility to meet the demands of real-world applications. We introduce the Tabular Auto-Regressive Generative Network (TabularARGN), a flexible framework designed to handle mixed-type, multivariate, and sequential datasets. By training on all possible conditional probabilities, TabularARGN supports advanced features such as fairness-aware generation, imputation, and conditional generation on any subset of columns. The framework achieves state-of-the-art synthetic data quality while significantly reducing training and inference times, making it ideal for large-scale datasets with diverse structures. Evaluated across established benchmarks, including realistic datasets with complex relationships, TabularARGN demonstrates its capability to synthesize high-quality data efficiently. By unifying flexibility and performance, this framework paves the way for practical synthetic data generation across industries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。