在部分共享潜变量下,实现多模态因果表征的可识别学习。
Identifiable Multimodal Causal Representation Learning under Partial Latent Sharing

- 基于非线性混合函数建模多模态数据生成机制。
- 在非参数假设下,证明了潜变量的逐分量可识别性。
- 适用于观测变量多于潜变量的欠定场景,适合高维数据建模。
因果表征学习(CRL)旨在从高维观测数据中发现有意义的潜变量及其对应的因果结构。尽管其重要性显著,但CRL的可识别性仍是关键属性,它确保能够恢复数据生成过程背后的机制,从而保证表征的可解释性和鲁棒性。证明CRL的可识别性本身极具挑战,本文进一步研究更复杂的多模态情形:观测数据具有部分共享的潜变量结构。每个模态通过非线性混合函数,由特定子集的因果潜变量生成。在灵活假设下,不施加任何关于潜变量的参数分布限制,我们建立了因果潜变量表示的逐分量可识别性保证。该结果还适用于欠定情形——每种模态的观测变量数量多于潜变量。为实例化理论分析,我们引入基于Wasserstein距离的模块以恢复部分共享的潜结构。由于其可微性,该模块可轻松集成到各类架构中,仅需最小修改。在合成与真实数据集上的大量实验验证了方法优于当前最先进方法。
原文摘要 · Abstract (English)
Causal representation learning (CRL) seeks to uncover meaningful latent variables and their corresponding causal structure from high-dimensional observational data. Although its significance, CRL identifiability remains a crucial property, as it ensures the recovery of the mechanisms behind the data generation process, and hence the interpretability and robustness of the representation. Proving identifiability in CRL is intrinsically difficult, and we address in this work an even more challenging setting: multimodality. We consider multimodal observed data with a latent partially shared structure. Each modality is generated, through non linear mixing functions, from a specific subset of causal latent variables. Under flexible assumptions and without imposing any parametric distribution on the latent variables, we establish component-wise identifiability guarantees for the causal latent representation. Our identifiability results, furthermore, apply to the undercomplete scenario where we have, for each modality, more observed than latent variables. To instantiate our theoretical analysis, we introduce a Wasserstein-based module to recover the partially shared latent structure. Due to its differentiability, the latter can be easily integrated into all types of architecture, only requiring minimal changes. Extensive experiments on synthetic and realistic datasets validate the superiority of our approach over SOTA methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。