arXiv:2509.04415cs.LG2025-09被引 3

该论文提出一种可解释的因果聚类方法,能从混合观测数据中自动发现异质性因果模式。

Interpretable Clustering with Adaptive Heterogeneous Causal Structure Learning in Mixed Observational Data

  • 通过双向迭代优化聚类与因果结构,自适应学习异质因果机制。
  • 在真实单细胞扰动数据上准确恢复出生物意义明确的因果机制。
  • 适合需揭示复杂系统内在因果差异的研究者,如生物医学领域。

理解因果异质性对生物学和医学等领域的科学发现至关重要。然而,现有方法缺乏因果意识,对异质性、混杂因素和观测约束建模不足,导致可解释性差,难以区分真实因果异质性与虚假关联。本文提出无监督框架HCL(可解释因果机制感知聚类与自适应异质因果结构学习),从混合类型观测数据中联合推断潜在聚类及其关联因果结构,无需时间顺序、环境标签、干预信息或其他先验知识。HCL通过引入等价表示,同时编码结构异质性与混杂关系,放松了同质性和充分性假设;并设计双向迭代策略交替优化因果聚类与结构学习,结合自监督正则化平衡跨簇普适性与特异性。这些组件共同促使模型收敛至可解释的异质因果模式。理论上,我们在较弱条件下证明了异质因果结构的可识别性。实验表明,HCL在聚类与结构学习任务上均表现更优,并在真实单细胞扰动数据中恢复出具有生物学意义的机制,验证其在发现可解释的机制级因果异质性方面的价值。

原文摘要 · Abstract (English)

Understanding causal heterogeneity is essential for scientific discovery in domains such as biology and medicine. However, existing methods lack causal awareness, with insufficient modeling of heterogeneity, confounding, and observational constraints, leading to poor interpretability and difficulty distinguishing true causal heterogeneity from spurious associations. We propose an unsupervised framework, HCL (Interpretable Causal Mechanism-Aware Clustering with Adaptive Heterogeneous Causal Structure Learning), that jointly infers latent clusters and their associated causal structures from mixed-type observational data without requiring temporal ordering, environment labels, interventions or other prior knowledge. HCL relaxes the homogeneity and sufficiency assumptions by introducing an equivalent representation that encodes both structural heterogeneity and confounding. It further develops a bi-directional iterative strategy to alternately refine causal clustering and structure learning, along with a self-supervised regularization that balance cross-cluster universality and specificity. Together, these components enable convergence toward interpretable, heterogeneous causal patterns. Theoretically, we show identifiability of heterogeneous causal structures under mild conditions. Empirically, HCL achieves superior performance in both clustering and structure learning tasks, and recovers biologically meaningful mechanisms in real-world single-cell perturbation data, demonstrating its utility for discovering interpretable, mechanism-level causal heterogeneity.

因果推断聚类分析单细胞数据可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。