通过因果差异网络识别生物干预靶点,助力药物与细胞工程研究。
Identifying biological perturbation targets through causal differential networks
- 构建观测与干预数据的因果图,学习其差异特征
- 在7个单细胞转录组数据集上优于基线方法
- 适合生物干预靶点预测,尤其适用于小样本场景
识别导致生物系统变化的变量,对药物靶点发现和细胞工程具有重要意义。给定一组观测数据和干预数据,目标是找出被干预的变量子集。直接应用因果发现算法面临挑战:数据通常包含数千变量,每种干预仅数十样本,且生物系统不满足经典因果假设。本文提出一种受因果启发的方法:首先从观测与干预数据中推断出含噪因果图;然后学习将两图差异及额外统计特征映射到被干预变量集合。两个模块在模拟与真实数据上联合监督训练,反映生物干预特性。该方法在七个单细胞转录组数据集上的扰动建模任务中持续优于基线。同时,在多种合成数据上,对软/硬干预靶点的预测也显著优于现有因果发现方法。
原文摘要 · Abstract (English)
Identifying variables responsible for changes to a biological system enables applications in drug target discovery and cell engineering. Given a pair of observational and interventional datasets, the goal is to isolate the subset of observed variables that were the targets of the intervention. Directly applying causal discovery algorithms is challenging: the data may contain thousands of variables with as few as tens of samples per intervention, and biological systems do not adhere to classical causality assumptions. We propose a causality-inspired approach to address this practical setting. First, we infer noisy causal graphs from the observational and interventional data. Then, we learn to map the differences between these graphs, along with additional statistical features, to sets of variables that were intervened upon. Both modules are jointly trained in a supervised framework, on simulated and real data that reflect the nature of biological interventions. This approach consistently outperforms baselines for perturbation modeling on seven single-cell transcriptomics datasets. We also demonstrate significant improvements over current causal discovery methods for predicting soft and hard intervention targets across a variety of synthetic data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。