综述扩散与流匹配模型在表格数据生成中的应用与挑战
Diffusion and Flow Matching Models for Tabular Data: A Survey
- 系统梳理扩散与流匹配模型在表格数据上的设计思路
- 涵盖生成、缺失值填补、异常检测等多任务应用
- 适合关注结构化数据生成的科研与工程人员
深度生成模型在图像、文本、音频和视频生成方面取得了快速发展,并逐步应用于结构化记录。然而,表格数据的生成仍面临诸多挑战:数据可能包含数值型与类别型属性、缺失值、敏感字段、类别不平衡、复杂特征依赖关系以及领域约束。基于GAN或VAE的早期方法虽取得一定成果,但存在训练不稳定、模式崩溃、多模态分布建模能力弱及混合类型特征处理脆弱等问题。扩散模型因噪声-去噪机制提供了灵活且稳定的复杂分布建模方式,已被用于表格数据合成、缺失值填补、可信数据生成与异常检测。流匹配则通过学习概率路径上的传输向量场,提供更直接的路径控制与采样效率。尽管进展显著,现有研究因任务目标、表示形式、优化目标、评估协议和领域假设各异,难以比较。本文首次专门综述扩散与流匹配模型在表格数据中的应用,覆盖2015年6月至2026年5月的研究工作,围绕数据工程挑战、任务、设计选择与评估维度进行组织,并讨论可扩展性、特征依赖建模、隐私、公平性、基准测试与约束感知生成等开放问题。相关更新维护于GitHub仓库。
原文摘要 · Abstract (English)
Deep generative models have made rapid progress in image, text, audio, and video generation, and are increasingly being applied to structured records. For tabular data, however, generative modeling remains difficult: a dataset may contain numerical and categorical attributes, missing values, sensitive fields, imbalanced categories, complex feature dependencies, and domain constraints. Earlier tabular data modeling methods based on GANs or VAEs have achieved useful results, but they can suffer from unstable training, mode collapse, weak modeling of multimodal distributions, and fragile handling of mixed-type features. Diffusion models have therefore attracted growing interest because their noising-and-denoising formulation provides a flexible and stable way to model complex data distributions, and has been adapted to tabular synthesis, missing-value imputation, trustworthy data generation, and anomaly detection. Flow matching offers a closely related route by learning transport vector fields along probability paths, often with more direct control over path design and sampling efficiency. Despite this progress, the literature on diffusion and flow matching models for tabular data remains difficult to compare because methods target different tasks and rely on different representations, objectives, evaluation protocols, and domain assumptions. To the best of our knowledge, this is the first survey dedicated specifically to diffusion and flow matching models for tabular data. We review work from June 2015 to May 2026, organize it around data-engineering challenges, tasks, design choices, and evaluation dimensions, and discuss open problems in scalability, feature dependency modeling, privacy, fairness, benchmarking, and constraint-aware generation. We maintain updates in a GitHub repository.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。