arXiv:2504.11547cs.AI2025-04被引 3

用因果图生成高质量分类数据,比高斯耦合和条件表生成对抗网络更优。

Probabilistic causal graphs as categorical data synthesizers: Do they do better than Gaussian Copulas and Conditional Tabular GANs?

  • 基于结构方程模型与贝叶斯网络构建因果关系,合成分类数据。
  • 在卡方检验、KL散度和总变差距离上优于高斯耦合与CTGAN。
  • 适合敏感数据研究,如残障人群服务可及性分析。

本研究探讨利用因果图模型生成高质量分类数据(如调查数据)的可行性。合成数据旨在用于模型训练并保护隐私,同时保留变量间关系。研究采用结构方程模型(SEM)结合贝叶斯网络(BN),基于残障人士服务可及性调查数据,建模人口统计、残疾类型、障碍类型及遭遇频率等变量间的因果关系与联合分布。对比高斯耦合与条件表生成对抗网络(CTGAN)等方法,所提方法在卡方检验、Kullback-Leibler散度和总变差距离(TVD)上表现更优。其中,贝叶斯网络模型在TVD指标上最高,表明其生成数据与原始数据高度一致。结果证实该方法能有效生成兼具统计真实性与隐私安全性的合成数据,适用于无障碍与残障研究等敏感领域。

原文摘要 · Abstract (English)

This study investigates the generation of high-quality synthetic categorical data, such as survey data, using causal graph models. Generating synthetic data aims not only to create a variety of data for training the models but also to preserve privacy while capturing relationships between the data. The research employs Structural Equation Modeling (SEM) followed by Bayesian Networks (BN). We used the categorical data that are based on the survey of accessibility to services for people with disabilities. We created both SEM and BN models to represent causal relationships and to capture joint distributions between variables. In our case studies, such variables include, in particular, demographics, types of disability, types of accessibility barriers and frequencies of encountering those barriers. The study compared the SEM-based BN method with alternative approaches, including the probabilistic Gaussian copula technique and generative models like the Conditional Tabular Generative Adversarial Network (CTGAN). The proposed method outperformed others in statistical metrics, including the Chi-square test, Kullback-Leibler divergence, and Total Variation Distance (TVD). In particular, the BN model demonstrated superior performance, achieving the highest TVD, indicating alignment with the original data. The Gaussian Copula ranked second, while CTGAN exhibited moderate performance. These analyses confirmed the ability of the SEM-based BN to produce synthetic data that maintain statistical and relational validity while maintaining confidentiality. This approach is particularly beneficial for research on sensitive data, such as accessibility and disability studies.

合成数据因果建模贝叶斯网络隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。