arXiv:2410.06163stat.MLcs.LG2024-10NeurIPS被引 7

通过正则化似然函数,让可微结构学习自动找到最稀疏的因果图。

Markov Equivalence and Consistency in Differentiable Structure Learning

  • 用正则化似然替代传统损失,确保优化收敛到最稀疏的马尔可夫等价模型。
  • 在无标识参数化条件下,仍能准确恢复马尔可夫等价类,且理论证明适用于一般模型。
  • 适合需要可微因果发现、尤其关注模型简洁性的研究者使用。

现有的可微有向无环图(DAG)结构学习方法依赖于强可识别性假设,以保证全局最小值能识别真实因果图。然而,实践中优化器可能利用损失函数中的不良特征。本文通过研究在多个全局最小值下的可微无环约束优化行为,解释并解决了这些问题。通过仔细正则化似然函数,即使在无标识参数化的情况下,也能识别出马尔可夫等价类中最稀疏的模型。我们首先详细分析高斯情形,证明适当正则化的似然可定义一个能识别最稀疏模型的评分函数;在忠实性假设下,还可恢复整个马尔可夫等价类。这些结论被推广至一般模型与似然函数,结果依然成立。实验证明,使用标准梯度优化器即可实现,为一般模型与损失下的可微结构学习铺平了道路。

原文摘要 · Abstract (English)

Existing approaches to differentiable structure learning of directed acyclic graphs (DAGs) rely on strong identifiability assumptions in order to guarantee that global minimizers of the acyclicity-constrained optimization problem identifies the true DAG. Moreover, it has been observed empirically that the optimizer may exploit undesirable artifacts in the loss function. We explain and remedy these issues by studying the behavior of differentiable acyclicity-constrained programs under general likelihoods with multiple global minimizers. By carefully regularizing the likelihood, it is possible to identify the sparsest model in the Markov equivalence class, even in the absence of an identifiable parametrization. We first study the Gaussian case in detail, showing how proper regularization of the likelihood defines a score that identifies the sparsest model. Assuming faithfulness, it also recovers the Markov equivalence class. These results are then generalized to general models and likelihoods, where the same claims hold. These theoretical results are validated empirically, showing how this can be done using standard gradient-based optimizers, thus paving the way for differentiable structure learning under general models and losses.

因果发现可微学习马尔可夫等价稀疏性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。