arXiv:2504.20687cs.LGstat.ML2025-04中稿 · , post peer-review…被引 4

用可解释AI分析生成数据缺陷,揭示隐藏的不真实模式

What's Wrong with Your Synthetic Tabular Data? Using Explainable AI to Evaluate Generative Models

  • 训练判别器识别真假数据,再用XAI技术分析差异原因
  • 发现传统指标忽略的异常依赖关系和缺失模式
  • 适合数据科学家诊断生成模型漏洞

评估合成表格数据质量极具挑战性,因其实质差异可能体现在多个层面。现有评估指标从统计距离到预测性能不一而足,常得出矛盾结论,且无法说明合成数据具体哪里出问题。为此,我们采用可解释人工智能(XAI)技术,对一个二分类判别器进行分析,该判别器用于区分真实与合成数据。尽管判别器能识别分布差异,但通过置换特征重要性、部分依赖图、Shapley值及反事实解释等方法,可解释性分析揭示了合成数据为何可被识别——暴露出不合理的依赖关系、异常模式或缺失结构。该方法提升了评估透明度,提供超越传统指标的深层洞察,有助于诊断并改进合成数据质量。我们在两个表格数据集和多种生成模型上验证了该方法,证明其能发现标准评估手段遗漏的问题。

原文摘要 · Abstract (English)

Evaluating synthetic tabular data is challenging, since they can differ from the real data in so many ways. There exist numerous metrics of synthetic data quality, ranging from statistical distances to predictive performance, often providing conflicting results. Moreover, they fail to explain or pinpoint the specific weaknesses in the synthetic data. To address this, we apply explainable AI (XAI) techniques to a binary detection classifier trained to distinguish real from synthetic data. While the classifier identifies distributional differences, XAI concepts such as feature importance and feature effects, analyzed through methods like permutation feature importance, partial dependence plots, Shapley values and counterfactual explanations, reveal why synthetic data are distinguishable, highlighting inconsistencies, unrealistic dependencies, or missing patterns. This interpretability increases transparency in synthetic data evaluation and provides deeper insights beyond conventional metrics, helping diagnose and improve synthetic data quality. We apply our approach to two tabular datasets and generative models, showing that it uncovers issues overlooked by standard evaluation techniques.

合成数据可解释AI数据质量表格生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。