FUSE模型显式分离数值与类别特征处理,提升混合类型表格数据生成质量。
FUSE: Feature-Wise Unified Specialization with Cross-Column Exchange for Mixed-Type Tabular Flow Matching

- 为数值和类别特征分别设计自适应混合模块,共享专用子网络
- 通过联合注意力机制保持跨列信息交换,提升生成数据一致性
- 在8个数据集上验证效果,兼顾分布保真度与下游任务性能
生成混合类型表格数据需联合建模多种特征分布及其复杂的跨列依赖关系。变分流匹配通过因子化分布处理不同终点,但将特征特异性处理与跨列交互隐含于共享主干中。本文提出特征级统一专业化与跨列交换(FUSE),显式分离这两项功能。FUSE为数值与类别特征分别应用自适应混合模块,使每个特征可组合共享的专用子网络,同时通过联合注意力机制保留所有列间的互信息。我们还分析了受限条件上下文带来的额外泛化风险,并以终点预测风险界定连续Wasserstein生成误差。在八个表格数据集上的全面实验表明,FUSE在分布保真度与下游任务效用指标上均表现出强且一致的性能。
原文摘要 · Abstract (English)
Generating mixed-type tabular data requires jointly modeling diverse feature distributions and their complex cross-column dependencies. Variational flow matching handles distinct endpoints via factorized distributions, yet leaves feature-specific processing and cross-column interactions implicit within a shared backbone. We introduce Feature-wise Unified Specialization with cross-column Exchange (FUSE) to explicitly separate these roles. FUSE applies separate adaptive mixture modules to numerical and categorical features, allowing each feature to combine shared specialized subnetworks, while joint attention preserves information exchange across all columns. We also characterize the excess population risk from restricted conditioning contexts and bound the continuous Wasserstein generation error by endpoint-prediction risk. Comprehensive experiments on eight tabular datasets demonstrate that FUSE achieves strong and consistent performance across distributional fidelity and downstream utility metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。