arXiv:2506.02084cs.LGstat.ML2025-06被引 3

用对抗训练优化因果模型,生成与真实数据分布一致的时间序列。

Adversarial Causal Tuning for Realistic Time-series Generation

  • 结合GAN和AutoML思想,自动搜索最优因果生成管道。
  • 在合成数据上避免过拟合,生成数据与真实数据无法区分。
  • 适合需要干预模拟和反事实推理的工业场景研究者。

我们解决从因果模型生成仿真但真实的时序数据问题,使其观测和干预分布与给定真实数据集一致(概率因果数字孪生)。非因果模型(如GAN)虽也追求数据真实性,但因果模型能模拟干预效果、优化决策、进行根因分析和反事实推理,具备根本优势。本文提出对抗因果调优(ACT)方法,输出最优因果模型并量化拟合优度。该方法借鉴生成对抗网络训练和AutoML思想,搜索最优因果流水线与检测真实与仿真数据分布差异的判别器,并采用置换检验程序惩罚模型复杂度。在真实、半合成和合成数据集上的大量实验表明:(a)使用多个优化判别器对选择最优因果模型和量化拟合优度至关重要;(b)ACT在合成数据上可选出最优模型且避免过拟合,生成数据与真实分布不可区分;(c)现有先进生成与因果模拟方法仍存在改进空间,真实时序数据生成仍是开放挑战。

原文摘要 · Abstract (English)

We address the problem of generating simulated, yet realistic, time-series data from a causal model with the same observational and interventional distributions as a given real dataset (probabilistic causal digital twin). While non-causal models (e.g., GANs) also strive to simulate realistic data, causal models are fundamentally more powerful, able to simulate the effect of interventions (what-if scenarios), optimize decisions, perform root-cause analysis, and counterfactual causal reasoning. We introduce the Adversarial Causal Tuning (ACT) methodology, which outputs the optimal causal model that fits the data, along with a quantification of the goodness-of-fit. The returned causal model can then be employed to simulate new data or to perform other causal reasoning tasks. ACT adopts ideas from Generative Adversarial Network training and AutoML to search for optimal causal pipelines and discriminators that detect deviations between the distributions of real and simulated data. It also adapts a permutation testing procedure from established causal tuning methods to penalize models for complexity. Through extensive experiments on real, semi-synthetic, and synthetic datasets, we show that (a) employing multiple optimized discriminators is paramount for selecting the optimal causal models and quantifying goodness-of-fit, (b) ACT selects the optimal causal model in synthetic datasets while avoiding overfitting, generating data indistinguishable from the true data distribution (c) all state-of-the-art generative and causal simulation methods, exhibit room for improvement in reproducing real data distributions; generating realistic temporal data is still an open research challenge.

因果建模时间序列生成对抗训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。