arXiv:2411.10982cs.LGstat.ME2024-11被引 3

用简单模型生成可解释的表格数据,适合做模型鲁棒性测试。

Towards a framework on tabular synthetic data generation: a minimalist approach: theory, use cases, and limitations

  • 用稀疏PCA编码+XGBoost解码生成合成数据
  • 在信贷评分数据上验证,效果优于传统扰动方法
  • 无需调参、全程可解释,适合重视透明度的场景

我们提出并研究了一种面向表格数据生成的极简方法。模型由最小化的无监督稀疏PCA编码器(可选聚类或对数变换处理非线性)和XGBoost解码器组成,后者在结构化数据回归与分类任务中表现领先。我们在多个低维模拟场景中对比了该方法与(变分)自编码器的差异,获得关键洞见。该框架应用于高维模拟信贷评分数据,其应用场景与真实金融业务相似。通过鲁棒性测试验证了其实用价值,结果表明该方法可作为原始数据或分位数扰动的替代方案用于模型稳健性评估。该方法简洁高效,保证全程可解释性,无需额外调参,具有独特优势。

原文摘要 · Abstract (English)

We propose and study a minimalist approach towards synthetic tabular data generation. The model consists of a minimalistic unsupervised SparsePCA encoder (with contingent clustering step or log transformation to handle nonlinearity) and XGboost decoder which is SOTA for structured data regression and classification tasks. We study and contrast the methodologies with (variational) autoencoders in several toy low dimensional scenarios to derive necessary intuitions. The framework is applied to high dimensional simulated credit scoring data which parallels real-life financial applications. We applied the method to robustness testing to demonstrate practical use cases. The case study result suggests that the method provides an alternative to raw and quantile perturbation for model robustness testing. We show that the method is simplistic, guarantees interpretability all the way through, does not require extra tuning and provide unique benefits.

表格生成可解释性鲁棒性测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。