融合多模态数据与因果推断,提升生物医学中的表征学习能力
Causal Structure and Representation Learning with Biomedical Applications
- 基于观测与干预数据构建因果结构学习框架
- 利用单细胞到生物体多尺度数据发现潜在因果变量
- 适用于生物医学中干预效果预测与最优扰动设计
海量数据收集有望深化对复杂现象的理解并支持更优决策。表征学习已成为深度学习应用的关键驱动力,它能在无需监督标注的情况下学习捕捉数据关键特征的隐空间。尽管表征学习在预测任务中表现卓越,但在因果任务(如预测干预后果)中常表现不佳。这促使表征学习与因果推断的结合成为必要。一个令人兴奋的机会来自多模态数据的日益丰富:包括观测与干预数据、成像与测序数据,覆盖单细胞、组织和生物体多个层级。本文提出一个统计与计算框架,旨在解决基础生物医学问题:如何有效利用观测与干预数据进行因果变量发现;如何通过系统多视角学习因果变量;以及如何设计最优干预方案。
原文摘要 · Abstract (English)
Massive data collection holds the promise of a better understanding of complex phenomena and, ultimately, better decisions. Representation learning has become a key driver of deep learning applications, as it allows learning latent spaces that capture important properties of the data without requiring any supervised annotations. Although representation learning has been hugely successful in predictive tasks, it can fail miserably in causal tasks including predicting the effect of a perturbation/intervention. This calls for a marriage between representation learning and causal inference. An exciting opportunity in this regard stems from the growing availability of multi-modal data (observational and perturbational, imaging-based and sequencing-based, at the single-cell level, tissue-level, and organism-level). We outline a statistical and computational framework for causal structure and representation learning motivated by fundamental biomedical questions: how to effectively use observational and perturbational data to perform causal discovery on observed causal variables; how to use multi-modal views of the system to learn causal variables; and how to design optimal perturbations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。