研究表格数据生成中潜在流模型的性能差异,给出实用配置建议。
Understanding Latent Flow Models for Tabular Data Synthesis: Targets, Paths, and Sampling
- 比较不同学习目标与采样路径对生成效果的影响
- 速度与后验匹配提升数据效用,梯度与噪声匹配降低泄露风险
- 提供可直接使用的默认配置,适配隐私与算力约束
合成表格数据可在受监管领域实现微观数据共享,但连续时间生成模型需在分析效用、泄露风险和计算成本间权衡。潜在空间流模型具有灵活性,但学习目标、概率路径与采样动态间的理论等价性,在有限步积分和明确计算预算下可能表现为不同行为。我们基于七个数据集,评估了速度、得分、噪声和后验匹配目标在最优传输(OT)与方差保持(VP)路径、常微分方程(ODE)与随机微分方程(SDE)采样、不同积分预算下的表现。主要贡献有三:(1) 学习目标显著决定效用-风险权衡,速度与后验匹配通常带来更高效用,而得分与噪声匹配更利于降低泄露风险;(2) 配置与采样选择影响性能,中点采样常提升分布保真度,OT路径通常可更早停止,优于VP路径,从而在固定预算或风险阈值下节省计算资源;(3) 将发现提炼为可操作默认设置与实用配置指南,支持发布前模型在隐私与资源限制下的选择。代码与补充材料见 https://github.com/rulnasution/tabular-latent-flow/。
原文摘要 · Abstract (English)
Synthetic tabular data enables microdata sharing in regulated domains, yet deploying continuous-time generative models requires balancing analytical utility, disclosure risk, and computational cost. Latent-space flow models are flexible, but theoretical equivalences across learning targets, probability paths, and sampling dynamics can translate into different behaviour under finite-step integration and explicit compute budgets. We present an empirical study of tabular latent flow models across seven datasets, evaluating velocity, score, noise, and posterior matching objectives under optimal transport (OT) and variance-preserving (VP) paths, ODE and SDE sampling, and varying integration budgets. Our contributions are threefold: (1) we show that the learning target largely determines the utility-risk operating regime, with velocity and posterior matching tending to yield higher utility, while score and noise matching tend to achieve lower disclosure risk; (2) we demonstrate that configuration and sampling choices shift performance, with midpoint often improving distributional fidelity and OT paths often tolerating earlier stopping than VP, enabling compute savings under fixed budgets or risk thresholds; and (3) we distil these findings into actionable defaults and practical configuration guidance to support pre-release model selection under disclosure risk and resource constraints. The code implementation and supplementary materials can be accessed in https://github.com/rulnasution/tabular-latent-flow/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。