同时聚类与学习异质因果结构,提升复杂系统中子群体依赖关系的发现精度。
A Unified Framework for Structure-Aware Clustering and Heterogeneous Causal Graph Learning

- 基于结构方程模型,联合优化聚类分配与各子群特异性因果图。
- 在真实数据上实现高真阳性率、低假阳性率,准确识别异质依赖结构。
- 适合研究具有隐藏子群体的复杂系统,如生物医学或社会行为数据。
在复杂的多变量系统中,变量间的相互作用由依赖结构定义,通常以有向无环图(DAG)编码。然而,这些依赖结构在不同个体间可能存在差异,忽略这种异质性会引入偏差并掩盖子群体特有的依赖关系。为此,我们提出基于ADMM的有向无环图依赖聚类方法(DAG-DC-ADMM),该框架基于结构方程模型(SEM),可联合学习样本聚类分配与各簇特异性依赖结构。通过平滑约束编码无环性,并引入组级截断Lasso融合惩罚(gTLP)以根据结构相似性聚类。该方法形成一个包含稀疏性、无环性及结构共识约束的非凸优化问题。我们采用增广拉格朗日法,并对凸差(DC)规划适配交替方向乘子法(ADMM)求解。对于某些图结构(如上三角邻接矩阵),算法保证收敛至KKT点。实验表明,该方法能以高真阳性率和低假发现率恢复子群特异性因果结构,从而在未知子群标签的情况下,稳健发现跨个体的异质依赖关系。
原文摘要 · Abstract (English)
In complex multivariate systems, interactions among variables are defined by dependency structures, often encoded as directed acyclic graphs ($\text{DAGs}$). However, dependency structures can vary across subjects, and ignoring this structural heterogeneity introduces bias and obscures subpopulation-specific dependencies. To address this, we propose Directed Acyclic Graph-based Dependency Clustering via Alternating Direction Method of Multipliers (DAG-DC-ADMM), a unified framework built upon Structural Equation Modeling (SEM) that jointly learns cluster assignments and cluster-specific dependency structures. We encode acyclicity via a smooth constraint and integrate a groupwise truncated Lasso fusion penalty (gTLP) to cluster subjects based on their structural similarity. This yields a nonconvex optimization problem that incorporates sparsity, acyclicity, and structural consensus constraints. We address the nonconvexity by using the augmented Lagrangian method and solve it with an adapted version of the Alternating Direction Method of Multipliers (ADMM) for difference-of-convex programs. For certain graph structures, such as upper triangular adjacency matrices, our algorithm is guaranteed to converge to a Karush-Kuhn-Tucker (KKT) point. Experiments demonstrate that our method recovers cluster-specific causal dependency structures with a high true positive rate and a low false discovery rate. This capability enables the robust discovery of heterogeneous dependencies across subjects where the subpopulation label is unknown.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。