arXiv:2607.08238cs.LGstat.ME2026-07

提出新方法在分组数据中同时学习全局因果结构与组内差异。

Structure Learning on Clustered Data

  • 将混合效应模型思想引入图结构学习,分离全局与局部效应。
  • 保证合并后的图无环,且在真实/合成数据上发现被忽略的依赖关系。
  • 适合医学、社会学等存在个体差异的因果推断场景。

近期算法进展使有向无环图(DAG)结构学习在因果发现中具备可扩展性,但现有方法假设总体同质,难以处理存在组内差异(如患者特异性效应)的分组数据。本文提出一种新方法,在估计全局结构的同时考虑局部集群级效应。核心思路是将经典混合模型中的固定效应与随机效应框架拓展至结构学习场景。为此,设计了一种可微的图耦合机制,确保固定与随机效应图的并集保持无环。计算上,提出一个保证收敛的一阶优化方法,并利用跨集群的高效批处理更新。统计上,证明了模型可识别性,并表明该方法能渐近恢复真实结构。在真实与合成数据上的实验显示,本方法检出了其他估计器遗漏的依赖关系,凸显其在分组结构学习中的价值。

原文摘要 · Abstract (English)

Recent algorithmic advances have made directed acyclic graph (DAG) structure learning scalable for causal discovery. Yet, the currently available techniques assume a completely homogeneous population, precluding their application to clustered data where cluster-specific variations (e.g., patient-specific effects) are common. We address this issue by introducing a new approach that estimates a global structure while accounting for local cluster-level effects. The key idea is to extend the fixed- and random-effects framework of classical mixed models to the structure learning setting. Towards this end, we present a differentiable graph coupling mechanism that guarantees the union of the fixed- and random-effects graphs remains acyclic. Computationally, we provide a provably convergent first-order method and leverage efficient batched updates across clusters. Statistically, we establish identifiability of the model and show that our approach recovers the true structure asymptotically. In experiments on real and synthetic data, our proposal detects dependencies missed by alternative estimators, underscoring its value for structure learning in clustered settings.

因果发现图学习分组数据混合效应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。