arXiv:2605.05524cs.LGcs.AI2026-05

从科学时间序列中自动发现可解释的模块化潜变量。

MOSAIC: Module Discovery via Sparse Additive Identifiable Causal Learning for Scientific Time Series

论文配图:MOSAIC: Module Discovery via Sparse Additive Identifiable Causal Learning for Scientific Time Series
图 1 · 摘自论文原文
  • 基于稀疏可加因果学习,通过时序变化识别潜变量。
  • 在多个真实科学数据集上恢复出一致的变量分组。
  • 适合需要挖掘潜因机制的科研人员使用。

因果表示学习(CRL)旨在恢复具有可识别性保证的潜变量,通常在合理假设下满足排列和分量重参数化不变性。然而,可识别性不等于可解释性:潜变量语义常需事后对齐已知真实因子。这一限制在科学时间序列中尤为突出,因底层机制未知,发现可解释结构是核心目标。相比之下,科学观测(如残基间距、气候指数或过程传感器)本身具有语义,对应命名物理量。这引发关键问题:能否将观测的可解释性传递到可识别的潜空间?本文提出MOSAIC(模块发现通过稀疏可加可识别因果学习),一种结合时序CRL可识别性与观测变量支持恢复的稀疏时序变分自编码器。MOSAIC通过分段条件时序变化识别潜变量,并通过可加解码器为每个潜变量恢复一个稀疏相关观测集合,实现模块级可解释性。我们证明,在一般光滑混合函数下,方差分析主效应支持可识别,并为可处理的稀疏可加变体提供有限样本恢复保证。实验表明,MOSAIC在RNA分子动力学、太阳风、厄尔尼诺-南方涛动气候、田纳西东部工业过程及合成托卡马克基准数据集中均能恢复域一致的变量组,支持科学时间序列中潜机制的可解释发现。

原文摘要 · Abstract (English)

Causal representation learning (CRL) seeks to recover latent variables with identifiability guarantees, typically up to permutation and component-wise reparameterization under appropriate assumptions. However, identifiability does not imply interpretability: latent semantics are typically assigned post hoc by alignment with known ground-truth factors. This limitation is particularly acute in scientific time series, where underlying mechanisms are unknown and discovering interpretable structure is a primary goal. In contrast, scientific observations (such as residue-pair distances, climate indices, or process sensors) are inherently semantic, as they correspond to named physical quantities. This raises a key question: can the interpretability of observations be transferred to the identifiable latent space? We propose MOSAIC (Module discovery via Sparse Additive Identifiable Causal learning), a sparse temporal VAE that integrates temporal CRL identifiability with support recovery over observed variables. MOSAIC identifies latent variables via regime-conditioned temporal variation, and recovers for each latent a sparse set of associated observations through an additive decoder, yielding module-level interpretability. We show that ANOVA main-effect supports are identifiable under general smooth mixing functions, and provide finite-sample recovery guarantees for a tractable sparse-additive variant. Empirically, MOSAIC recovers domain-consistent variable groups across RNA molecular dynamics, solar wind, ENSO climate, the Tennessee Eastman process, and a synthetic tokamak benchmark, enabling interpretable discovery of latent mechanisms in scientific time series.

因果学习时间序列可解释性科学建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。