用对数坐标生成混合型表格数据,提升稀有类别建模效果
Logit-Coordinate Generative Models for Mixed Continuous-Categorical Tabular Data

- 将类别变量转为对数坐标,与数值变量统一处理
- 在四个真实数据集上,对数流匹配优于或等同于独热编码
- 特别适合稀有类别不平衡场景,如客户流失预测
混合连续-类别数据对连续生成模型构成表征难题:流匹配和高斯扩散在欧氏空间中运行,而类别分布位于概率单纯形上且可能严重失衡。本文提出对数坐标框架,将类别变量编码为平滑的自然参数,并与变换后的数值变量结合,统一构建对数流匹配与对数扩散模型。引入混合分布差异度量,分离类别边缘误差与条件连续Wasserstein误差,推导出向量场或漂移误差与解码后混合分布误差间的稳定性界与平衡感知的非参数率。受控模拟显示,缩放对数坐标在严重稀有单元失衡下优于或等同于独热编码;在四个真实数据集、每数据集十次划分的实验中,对数流匹配在三个数据集上提升主要分布指标,与客户流失数据集(Churn2)表现相当;块条件对数流匹配始终优于平坦模型;对数扩散总体优于或等同于独热扩散。
原文摘要 · Abstract (English)
Mixed continuous--categorical data pose a representation problem for continuous generative models. Flow Matching and Gaussian diffusion operate in Euclidean spaces, whereas categorical laws lie on probability simplices and may be highly imbalanced. We study a logit-coordinate framework that encodes categorical variables as smoothed natural parameters and combines them with transformed numerical variables. This yields common formulations of Logit Flow Matching and Logit Diffusion. We introduce a mixed-distribution discrepancy separating categorical marginal error from conditional continuous Wasserstein error, and derive stability bounds and imbalance-aware nonparametric rates linking vector-field or drift error to decoded mixed-distribution error. Controlled simulations show that scaled-logit coordinates improve or match one-hot coordinates, especially under severe rare-cell imbalance. Across four real-data benchmarks and ten splits per dataset, Logit FM improves the primary distributional metrics on three datasets and is comparable on Churn2; Block-Conditional Logit FM consistently improves the flat model; and Logit Diffusion generally improves over or matches One-Hot Diffusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。