arXiv:2609.03705cs.LG2026-09

提出可联邦化且抗对称噪声的因果发现算法,解决隐私与精度矛盾。

Federated Causal Discovery via Regression-Directed Cumulants

论文配图:Federated Causal Discovery via Regression-Directed Cumulants
图 1 · 摘自论文原文
  • 用高阶累积量设计单轮通信联邦算法,支持多种数据划分模式。
  • 在真实样本规模下,算法按变量方差梯度排序而非对称性,具稳定性能。
  • 支持细粒度数据删除,适合需合规隐私保护的医疗、金融场景。

本文研究在联邦环境下使用线性非高斯无环模型(LiNGAM)进行因果发现。这类模型可突破马尔可夫等价性限制,但实际中数据稀缺,集中数据受GDPR等法规约束,难以实现。联邦环境为平衡隐私与准确性提供可行路径。然而,标准集中式估计器DirectLiNGAM无法直接联邦化。高阶累积量张量仅依赖变量联合分布,且在独立样本组间可加,可在水平、垂直及混合划分下实现单轮通信。但现有方法FedISHC在近对称噪声下失效。为此,我们提出FedRCD族算法,包含三个变体,权衡通信轮次与代数噪声鲁棒性;其中两个为集中式高阶累积量(HC)与HC-LiNGAM的精确联邦对应,单轮版本还可实现任意粒度(从单个样本到整个客户端)的精确数据删除。数值实验表明,在典型实际部署样本量下,所有基于累积量的联邦方法并未按总体不对称性排序,而是依据由有向无环图(DAG)路径诱导的方差梯度排序——即累积量版的varsortability。边际标准化使所有累积量方法退化为近乎随机排序,而尺度不变的DirectLiNGAM虽不满足该协议,却不受影响。

原文摘要 · Abstract (English)

In this paper we study linear non-Gaussian acyclic models (LiNGAM) when used in federated environments. These causal models allow one to go beyond Markov equivalence. However, in many domains data are scarce, and increasing the sample size by centralising data from different clients is not advisable due to regulations such as the GDPR. The federated environment offers an attractive option to balance privacy and causal discovery accuracy. Unfortunately, the standard centralised estimator in the LiNGAM setting, i.e., DirectLiNGAM, cannot be straightforwardly federated. Higher-order cumulant tensors offer a way around this obstacle: they depend only on the joint distribution of the variables involved and add exactly across independent sample groups, so a single communication round suffices in horizontal, vertical, and hybrid partitions. However, FedISHC, i.e., the current federated method along these lines, breaks down under near-symmetric noise. To overcome the above limitation, we introduce the FedRCD family of causal discovery algorithms, and investigate three variants that trade off communication rounds against algebraic noise; two of them are exact federated counterparts of the centralised high-order cumulant (HC) and HC-LiNGAM algorithms, and the single-round variants further effectively support exact unlearning at any granularity, from a single observation to a whole client. Numerical experiments show that at sample sizes typical of real deployments, the entire cumulant-based federated family does not actually rank variables by the population asymmetry that the scores encode at zero. It ranks them by a variance ladder induced by the DAG along its directed paths, the cumulant counterpart of varsortability. Marginal standardisation collapses every cumulant method to near-random ordering, while scale-invariant DirectLiNGAM, not federable under this protocol, is unaffected.

联邦学习因果发现隐私保护高阶统计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。