提出可识别的连续时间因果表示学习框架,解决真实世界动态数据建模难题。
Causal Representation Meets Stochastic Modeling under Generic Geometry
- 基于参数空间几何分析,实现连续时间随机点过程的可识别因果表征
- 构建MUTATE模型,通过时变转移模块有效捕捉动态演化规律
- 适用于基因突变积累、神经元放电等科学问题,具备强解释性
从观测中学习有意义的因果表示已成为推动机器学习应用和气候科学、生物学、物理学等领域科学发现的关键任务。该过程需从低层观测中解耦高层潜在变量及其因果关系。以往工作在实现可识别性方面多集中于独立同分布或离散时间潜变量过程,但许多现实场景需要识别连续时间随机过程(如多变量点过程)。为此,本文提出针对连续时间潜随机点过程的可识别因果表示学习方法。通过分析参数空间的几何结构研究其可识别性,并开发MUTATE——一种具有时变转移模块的可识别变分自编码器框架,用于推断随机动力学。在模拟与实证研究中,MUTATE能有效回答科学问题,例如基因组中突变的累积机制以及神经元对时变动态刺激的放电触发机理。
原文摘要 · Abstract (English)
Learning meaningful causal representations from observations has emerged as a crucial task for facilitating machine learning applications and driving scientific discoveries in fields such as climate science, biology, and physics. This process involves disentangling high-level latent variables and their causal relationships from low-level observations. Previous work in this area that achieves identifiability typically focuses on cases where the observations are either i.i.d. or follow a latent discrete-time process. Nevertheless, many real-world settings require identifying latent variables that are continuous-time stochastic processes (e.g., multivariate point processes). To this end, we develop identifiable causal representation learning for continuous-time latent stochastic point processes. We study its identifiability by analyzing the geometry of the parameter space. Furthermore, we develop MUTATE, an identifiable variational autoencoder framework with a time-adaptive transition module to infer stochastic dynamics. Across simulated and empirical studies, we find that MUTATE can effectively answer scientific questions, such as the accumulation of mutations in genomics and the mechanisms driving neuron spike triggers in response to time-varying dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。