arXiv:2505.02238cs.LG2025-05被引 3

联邦因果推断让多机构合作研究治疗效果,不碰原始数据也能算得准。

Federated Causal Inference in Healthcare: Methods, Challenges, and Applications

  • 按加权与优化两类方法分类,提出统一分析框架
  • 证明FedProx正则化在异质数据下接近最优偏差-方差平衡
  • 适合医疗数据隐私敏感场景的科研人员和临床研究者

联邦因果推断可在不共享个体数据的前提下实现多中心治疗效果估计,为真实世界证据生成提供隐私保护方案。然而,各机构间协变量、处理和结果分布的差异带来了显著偏差与效率挑战。本文系统综述并理论分析了二元/连续及生存时间结果的联邦因果效应估计方法。将现有方法分为基于权重与基于优化的框架,并进一步讨论个性化模型、点对点通信与模型分解等扩展。针对生存时间结果,分析了联邦Cox与Aalen-Johansen模型,在异质性条件下推导其渐近偏差与方差。分析表明,相比简单平均与元分析,FedProx式正则化可实现近似最优的偏差-方差权衡。本文还回顾相关软件工具,并展望可扩展、公平且可信的分布式医疗系统中联邦因果推断的发展机遇与挑战。

原文摘要 · Abstract (English)

Federated causal inference enables multi-site treatment effect estimation without sharing individual-level data, offering a privacy-preserving solution for real-world evidence generation. However, data heterogeneity across sites, manifested in differences in covariate, treatment, and outcome, poses significant challenges for unbiased and efficient estimation. In this paper, we present a comprehensive review and theoretical analysis of federated causal effect estimation across both binary/continuous and time-to-event outcomes. We classify existing methods into weight-based strategies and optimization-based frameworks and further discuss extensions including personalized models, peer-to-peer communication, and model decomposition. For time-to-event outcomes, we examine federated Cox and Aalen-Johansen models, deriving asymptotic bias and variance under heterogeneity. Our analysis reveals that FedProx-style regularization achieves near-optimal bias-variance trade-offs compared to naive averaging and meta-analysis. We review related software tools and conclude by outlining opportunities, challenges, and future directions for scalable, fair, and trustworthy federated causal inference in distributed healthcare systems.

联邦学习因果推断医疗数据隐私计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。