arXiv:2411.06518cs.LGq-bio.QM2024-11ICLR被引 16

从多模态生物医学数据中学习可解释的因果表示,提升模型可理解性。

Causal Representation Learning from Multimodal Biomedical Observations

  • 基于非参数潜变量建模,利用模态间结构稀疏性实现因果关系识别
  • 在真实人类表型数据上验证了与已有医学研究一致的结果
  • 适合需要机制可解释性的生物医学分析场景

多模态生物医学数据(如人类表型研究)蕴含丰富的生理机制信息。然而,现有机器学习模型常缺乏可解释性和可辨识性保证,制约其在生物医学研究中的应用。近年来因果表示学习虽在识别可解释潜变量方面展现潜力,但多数方法依赖严格参数假设或仅能获得粗粒度辨识结果,难以满足生物医学对机制细节的需求。本文提出一种灵活的多模态数据辨识条件与方法框架。理论上,采用非参数潜分布建模,允许跨模态潜在因果关系,并建立各潜成分的可辨识性保证,扩展了先前子空间辨识结果。关键理论贡献在于揭示模态间因果连接的结构稀疏性,该特性在大量生物系统中自然存在。实证上,构建实用化方法框架,在数值与合成数据上充分验证有效性;在真实人类表型数据上的结果与已有生物医学研究一致,证实了理论与方法的合理性。

原文摘要 · Abstract (English)

Prevalent in biomedical applications (e.g., human phenotype research), multimodal datasets can provide valuable insights into the underlying physiological mechanisms. However, current machine learning (ML) models designed to analyze these datasets often lack interpretability and identifiability guarantees, which are essential for biomedical research. Recent advances in causal representation learning have shown promise in identifying interpretable latent causal variables with formal theoretical guarantees. Unfortunately, most current work on multimodal distributions either relies on restrictive parametric assumptions or yields only coarse identification results, limiting their applicability to biomedical research that favors a detailed understanding of the mechanisms. In this work, we aim to develop flexible identification conditions for multimodal data and principled methods to facilitate the understanding of biomedical datasets. Theoretically, we consider a nonparametric latent distribution (c.f., parametric assumptions in previous work) that allows for causal relationships across potentially different modalities. We establish identifiability guarantees for each latent component, extending the subspace identification results from previous work. Our key theoretical contribution is the structural sparsity of causal connections between modalities, which, as we will discuss, is natural for a large collection of biomedical systems. Empirically, we present a practical framework to instantiate our theoretical insights. We demonstrate the effectiveness of our approach through extensive experiments on both numerical and synthetic datasets. Results on a real-world human phenotype dataset are consistent with established biomedical research, validating our theoretical and methodological framework.

因果学习多模态生物医学可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。