针对多源异构数据,用非线性因果核聚类识别不同群体的差异因果关系。
Causal Learning for Heterogeneous Subgroups Based on Nonlinear Causal Kernel Clustering
- 通过无偏估计的u中心样本映射函数捕捉非线性因果差异
- 实验显示能有效识别异质子群并降低预测误差
- 适合需要细分因果关系的跨区域或跨时间研究
由于来自不同环境的多源异构数据带来的挑战,特征间的因果关系会受时间跨度、地区或策略差异的影响而变化。单一因果模型难以准确刻画所有观测数据中的复杂因果结构,这在因果学习中至关重要。为此,本文提出非线性因果核聚类方法,用于异质子群的因果学习,突出不同子群间因果关系的差异。核心在于构建具有无偏估计性质的u中心样本映射函数,评估各样本间潜在的非线性因果关系差异,并由因果可识别性理论支持。实验结果表明,该方法在识别异质子群和提升因果学习性能方面表现良好,显著降低了预测误差。
原文摘要 · Abstract (English)
Due to the challenge posed by multi-source and heterogeneous data collected from diverse environments, causal relationships among features can exhibit variations influenced by different time spans, regions, or strategies. This diversity makes a single causal model inadequate for accurately representing complex causal relationships in all observational data, a crucial consideration in causal learning. To address this challenge, the nonlinear Causal Kernel Clustering method is introduced for heterogeneous subgroup causal learning, highlighting variations in causal relationships across diverse subgroups. The main component for clustering heterogeneous subgroups lies in the construction of the $u$-centered sample mapping function with the property of unbiased estimation, which assesses the differences in potential nonlinear causal relationships in various samples and supported by causal identifiability theory. Experimental results indicate that the method performs well in identifying heterogeneous subgroups and enhancing causal learning, leading to a reduction in prediction error.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。