arXiv:2502.20925stat.MLcs.LG2025-02KDD被引 1

用神经网络自动测试变量间条件独立性,高效且可迁移。

Amortized Conditional Independence Testing

  • 设计基于Transformer的ACID模型,通过合成数据训练实现条件独立性检测
  • 在多种数据上表现超越现有方法,支持不同样本量和非线性关系
  • 训练后推理极快,可低开销适配新领域,适合需要快速分析的场景

在统计与机器学习中,检验数据中的条件独立结构是一项基础而关键的任务,广泛应用于因果发现。现有方法依赖显式检验统计量来量化条件依赖程度,但设计困难,且难以数据驱动地利用先验知识。本文提出一种全新方法:将条件独立性测试过程进行摊销,并设计ACID——一种基于Transformer的新型神经网络架构,用于学习条件独立性检验。ACID可在合成数据上以监督学习方式训练,训练完成后可直接应用于同类型数据,或通过微调低计算成本适配新领域。在合成与真实数据上的大量实验表明,ACID在多个指标上持续达到最先进性能,对未见样本量、维度及非线性具有强泛化能力,且推理时间极短。

原文摘要 · Abstract (English)

Testing for the conditional independence structure in data is a fundamental and critical task in statistics and machine learning, which finds natural applications in causal discovery - a highly relevant problem to many scientific disciplines. Existing methods seek to design explicit test statistics that quantify the degree of conditional dependence, which is highly challenging yet cannot capture nor utilize prior knowledge in a data-driven manner. In this study, an entirely new approach is introduced, where we instead propose to amortize conditional independence testing and devise ACID - a novel transformer-based neural network architecture that learns to test for conditional independence. ACID can be trained on synthetic data in a supervised learning fashion, and the learned model can then be applied to any dataset of similar natures or adapted to new domains by fine-tuning with a negligible computational cost. Our extensive empirical evaluations on both synthetic and real data reveal that ACID consistently achieves state-of-the-art performance against existing baselines under multiple metrics, and is able to generalize robustly to unseen sample sizes, dimensionalities, as well as non-linearities with a remarkably low inference time.

条件独立因果发现神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。