arXiv:2511.17590cs.LGcs.AI2025-11被引 3

提出新指标评估合成表格数据的语义保真度,捕捉传统方法遗漏的特征重要性偏差。

SHAP Distance: An Explainability-Aware Metric for Evaluating the Semantic Fidelity of Synthetic Tabular Data

  • 基于真实与合成数据训练模型的SHAP全局归因向量间余弦距离
  • 能检测出KL散度和预测准确率未发现的特征重要性偏移与尾部效应
  • 适合关注数据语义一致性的医疗、金融等领域的合成数据质量评估

合成表格数据广泛应用于医疗、企业运营和客户分析等领域,其评估需兼顾隐私保护与实用性。现有方法多聚焦分布相似性(如KL散度)或预测性能(如TSTR准确率),但无法衡量语义保真度——即在合成数据上训练的模型是否遵循与真实数据一致的推理模式。为此,本文提出SHAP Distance,一种可解释性感知的新指标,定义为在真实与合成数据上训练的分类器所得全局SHAP归因向量间的余弦距离。通过对包含生理特征的临床健康记录、异构尺度的企业发票交易及混合类别-数值属性的电信流失日志等数据集的分析,结果表明:该指标能可靠识别出标准统计与预测指标所忽略的语义差异。尤其发现,其可捕捉特征重要性转移与被低估的尾部效应,而这些正是KL散度与TSTR准确率未能检测到的。本研究将SHAP Distance定位为审计合成表格数据语义保真度的实用且具有判别力的工具,并为未来基准测试流程中集成基于归因的评估提供实践指南。

原文摘要 · Abstract (English)

Synthetic tabular data, which are widely used in domains such as healthcare, enterprise operations, and customer analytics, are increasingly evaluated to ensure that they preserve both privacy and utility. While existing evaluation practices typically focus on distributional similarity (e.g., the Kullback-Leibler divergence) or predictive performance (e.g., Train-on-Synthetic-Test-on-Real (TSTR) accuracy), these approaches fail to assess semantic fidelity, that is, whether models trained on synthetic data follow reasoning patterns consistent with those trained on real data. To address this gap, we introduce the SHapley Additive exPlanations (SHAP) Distance, a novel explainability-aware metric that is defined as the cosine distance between the global SHAP attribution vectors derived from classifiers trained on real versus synthetic datasets. By analyzing datasets that span clinical health records with physiological features, enterprise invoice transactions with heterogeneous scales, and telecom churn logs with mixed categorical-numerical attributes, we demonstrate that the SHAP Distance reliably identifies semantic discrepancies that are overlooked by standard statistical and predictive measures. In particular, our results show that the SHAP Distance captures feature importance shifts and underrepresented tail effects that the Kullback-Leibler divergence and Train-on-Synthetic-Test-on-Real accuracy fail to detect. This study positions the SHAP Distance as a practical and discriminative tool for auditing the semantic fidelity of synthetic tabular data, and offers practical guidelines for integrating attribution-based evaluation into future benchmarking pipelines.

合成数据可解释性语义保真表格数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。