让基因扰动预测模型可解释,每个潜变量对应生物通路活动。
Interpretable Causal Representation Learning for Biological Data in the Pathway Space
- 基于差异VAE构建可解释的因果表示学习模型
- 在未见干预组合上保持与原始模型相当的预测性能
- 适合关注生物机制可解释性的药物研发研究者
预测基因和药物扰动对细胞功能的影响对于理解基因功能和药物效应至关重要,有助于改进疗法。因果表示学习(CRL)因其能识别驱动生物系统的潜在因果因子而成为最有前景的方法之一。然而,现有CRL方法难以将其潜变量与已知生物过程对齐,导致模型不可解释。为此,我们提出SENA-discrepancy-VAE,基于近期提出的discrepancy-VAE方法,生成每个潜因子可解释为(学习到的)生物过程活动水平线性组合的表示。我们设计了SENA-δ编码器,高效计算并映射生物过程的活动水平至潜因子里。结果表明,SENA-discrepancy-VAE在未见干预组合上的预测性能与原始非可解释模型相当,同时推断出具有生物学意义的因果潜因子。
原文摘要 · Abstract (English)
Predicting the impact of genomic and drug perturbations in cellular function is crucial for understanding gene functions and drug effects, ultimately leading to improved therapies. To this end, Causal Representation Learning (CRL) constitutes one of the most promising approaches, as it aims to identify the latent factors that causally govern biological systems, thus facilitating the prediction of the effect of unseen perturbations. Yet, current CRL methods fail in reconciling their principled latent representations with known biological processes, leading to models that are not interpretable. To address this major issue, we present SENA-discrepancy-VAE, a model based on the recently proposed CRL method discrepancy-VAE, that produces representations where each latent factor can be interpreted as the (linear) combination of the activity of a (learned) set of biological processes. To this extent, we present an encoder, SENA-δ, that efficiently compute and map biological processes' activity levels to the latent causal factors. We show that SENA-discrepancy-VAE achieves predictive performances on unseen combinations of interventions that are comparable with its original, non-interpretable counterpart, while inferring causal latent factors that are biologically meaningful.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。