提出EF-VFM方法,实现混合类型表格数据的高效生成。
Exponential Family Variational Flow Matching for Tabular Data Generation
- 用指数族分布统一建模连续与离散特征,实现混合数据生成。
- 基于矩匹配构建数据驱动目标,学习跨类型概率路径。
- 在多个基准上超越现有方法,适合真实场景表格数据生成。
尽管去噪扩散和流匹配在生成建模中取得显著进展,但其在表格数据上的应用仍受限,而表格数据在现实场景中广泛存在。为此,我们提出TabbyFlow,一种用于表格数据生成的变分流匹配(VFM)方法。为将VFM应用于包含连续与离散特征的数据,我们引入指数族变分流匹配(EF-VFM),通过广义指数族分布表示异构数据类型,从而得到基于矩匹配的高效、数据驱动的目标函数,实现对混合连续与离散变量的概率路径的合理建模。我们还建立了变分流匹配与基于Bregman散度的广义流匹配目标之间的联系。在多个表格数据基准上的评估表明,该方法性能优于现有基线。
原文摘要 · Abstract (English)
While denoising diffusion and flow matching have driven major advances in generative modeling, their application to tabular data remains limited, despite its ubiquity in real-world applications. To this end, we develop TabbyFlow, a variational Flow Matching (VFM) method for tabular data generation. To apply VFM to data with mixed continuous and discrete features, we introduce Exponential Family Variational Flow Matching (EF-VFM), which represents heterogeneous data types using a general exponential family distribution. We hereby obtain an efficient, data-driven objective based on moment matching, enabling principled learning of probability paths over mixed continuous and discrete variables. We also establish a connection between variational flow matching and generalized flow matching objectives based on Bregman divergences. Evaluation on tabular data benchmarks demonstrates state-of-the-art performance compared to baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。