针对少数类数据生成难题,提出结构化隐空间与自适应采样策略。
CTTVAE: Latent Space Structuring for Conditional Tabular Data Generation on Imbalanced Datasets
- 用类别感知三元组损失重构隐空间,增强类内紧凑与类间分离。
- 在6个真实数据集上,对少数类的下游性能显著优于原始数据训练模型。
- 适合医疗、反欺诈等少数类关键场景,兼顾生成质量与实际应用价值。
在严重类别不平衡的数据上生成合成表格数据对高影响事件驱动的领域至关重要。然而,多数生成模型要么忽略少数类,要么无法生成对下游学习有用的样本。我们提出CTTVAE,一种基于条件Transformer的表格变分自编码器,包含两个互补机制:(i) 类别感知三元组边界损失,重构隐空间以实现更清晰的类内紧凑性和类间分离;(ii) 基于采样的训练策略,自适应增加对少数类的暴露。二者结合形成CTTVAE+TBS框架,在不破坏训练稳定性的情况下持续生成更具代表性且符合下游任务需求的样本。在六个真实世界基准测试中,该方法在少数类上的下游性能最强,常超越在原始不平衡数据上训练的模型,同时保持良好保真度,并缩小了基于插值采样方法与深度生成方法之间的差距。消融实验进一步验证了隐空间结构化与针对性采样均带来增益。通过明确提升稀有类别表现,CTTVAE+TBS为条件表格数据生成提供了一种稳健且可解释的解决方案,适用于医疗、欺诈检测、预测性维护等少数类决策至关重要的行业。
原文摘要 · Abstract (English)
Generating synthetic tabular data under severe class imbalance is essential for domains where rare but high-impact events drive decision-making. However, most generative models either overlook minority groups or fail to produce samples that are useful for downstream learning. We introduce CTTVAE, a Conditional Transformer-based Tabular Variational Autoencoder equipped with two complementary mechanisms: (i) a class-aware triplet margin loss that restructures the latent space for sharper intra-class compactness and inter-class separation, and (ii) a training-by-sampling strategy that adaptively increases exposure to underrepresented groups. Together, these components form CTTVAE+TBS, a framework that consistently yields more representative and utility-aligned samples without destabilizing training. Across six real-world benchmarks, CTTVAE+TBS achieves the strongest downstream utility on minority classes, often surpassing models trained on the original imbalanced data while maintaining competitive fidelity and bridging the gap for privacy for interpolation-based sampling methods and deep generative methods. Ablation studies further confirm that both latent structuring and targeted sampling contribute to these gains. By explicitly prioritizing downstream performance in rare categories, CTTVAE+TBS provides a robust and interpretable solution for conditional tabular data generation, with direct applicability to industries such as healthcare, fraud detection, and predictive maintenance where even small gains in minority cases can be critical.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。