arXiv:2509.19058cs.AI2025-09被引 3

让可观察变量辅助识别数据背后的因果因子,突破传统方法限制。

Towards Causal Representation Learning with Observable Sources as Auxiliaries

  • 用可观察变量作为条件变量,提升潜在因子的可识别性。
  • 通过保体积编码器,实现潜变量在子空间变换下的完全识别。
  • 提出变量选择策略,基于因果图优化辅助变量选取,适合因果建模研究者。

因果表示学习旨在通过混合函数恢复生成观测数据的潜在因素。现有方法通常依赖已知辅助变量的条件独立性假设来实现可识别性,但以往框架将辅助变量限制为混合函数外部的变量。然而,在某些情况下,系统驱动的潜在因子可直接从数据中观测或提取,可能有助于识别。本文提出一种新框架:将可观测源作为辅助变量,作为有效的条件变量。主要结果表明,利用保体积编码器,可在子空间变换和排列意义下完全识别所有潜变量。当存在多个已知辅助变量时,我们设计了一种变量选择方案,根据潜在因果图知识最大化潜因子的可恢复性。最后,我们在合成图和图像数据上验证了该框架的有效性,拓展了当前方法的边界。

原文摘要 · Abstract (English)

Causal representation learning seeks to recover latent factors that generate observational data through a mixing function. Needing assumptions on latent structures or relationships to achieve identifiability in general, prior works often build upon conditional independence given known auxiliary variables. However, prior frameworks limit the scope of auxiliary variables to be external to the mixing function. Yet, in some cases, system-driving latent factors can be easily observed or extracted from data, possibly facilitating identification. In this paper, we introduce a framework of observable sources being auxiliaries, serving as effective conditioning variables. Our main results show that one can identify entire latent variables up to subspace-wise transformations and permutations using volume-preserving encoders. Moreover, when multiple known auxiliary variables are available, we offer a variable-selection scheme to choose those that maximize recoverability of the latent factors given knowledge of the latent causal graph. Finally, we demonstrate the effectiveness of our framework through experiments on synthetic graph and image data, thereby extending the boundaries of current approaches.

因果学习表示学习可识别性辅助变量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。