arXiv:2602.19782cs.LG2026-02

用跨环境不变性学习,解决遗传工具变量的混杂偏差问题。

Addressing Instrument-Outcome Confounding in Mendelian Randomization through Representation Learning

  • 通过多环境数据学习遗传工具的不变表示,剥离混杂因素影响。
  • 在模拟和真实数据中验证,显著提升因果效应估计的准确性。
  • 适合做遗传学因果推断、生物统计或医学大数据分析的研究者。

孟德尔随机化(Mendelian Randomization, MR)是一种用于估算因果效应的流行观察性研究方法,旨在缓解未观测混杂因素的影响。然而,其核心假设——工具变量与未观测混杂因子独立——常因人群分层或择偶偏好而被破坏。利用日益丰富的多环境数据,我们提出一种表示学习框架,通过挖掘跨环境不变性来恢复遗传工具变量的潜在外生成分。我们在多种混合机制下提供了理论保障,证明可识别这些潜在工具变量,并通过来自All of Us Research Hub的半真实实验和模拟验证了该方法的有效性。

原文摘要 · Abstract (English)

Mendelian Randomization (MR) is a prominent observational epidemiological research method designed to address unobserved confounding when estimating causal effects. However, core assumptions -- particularly the independence between instruments and unobserved confounders -- are often violated due to population stratification or assortative mating. Leveraging the increasing availability of multi-environment data, we propose a representation learning framework that exploits cross-environment invariance to recover latent exogenous components of genetic instruments. We provide theoretical guarantees for identifying these latent instruments under various mixing mechanisms and demonstrate the effectiveness of our approach through simulations and semi-synthetic experiments using data from the All of Us Research Hub.

因果推断遗传学表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。