arXiv:2412.11511cs.LGstat.ML2024-12ICLR被引 6

多来源医疗数据合起来更准地评估药物效果

Constructing Confidence Intervals for Average Treatment Effects from Multiple Datasets

  • 用预测增强推断方法合并多个医院数据
  • 相比传统方法,置信区间更窄,误差更小
  • 适合医学研究中真实世界数据整合场景

从患者记录构建平均治疗效应(ATE)的置信区间(CI)对评估药物疗效与安全性至关重要。然而,患者数据通常来自不同医院,如何有效整合多个观察性数据集成为关键问题。本文提出一种新方法,可在少假设条件下估计多个观察性数据集中的ATE并提供有效置信区间。核心思想是利用预测增强推断,从而‘收缩’置信区间,实现比朴素方法更精确的不确定性量化。我们证明了该方法的无偏性及置信区间的有效性,并通过多种数值实验验证理论结果。最后,方法扩展至结合实验与观察数据集的情形。

原文摘要 · Abstract (English)

Constructing confidence intervals (CIs) for the average treatment effect (ATE) from patient records is crucial to assess the effectiveness and safety of drugs. However, patient records typically come from different hospitals, thus raising the question of how multiple observational datasets can be effectively combined for this purpose. In our paper, we propose a new method that estimates the ATE from multiple observational datasets and provides valid CIs. Our method makes little assumptions about the observational datasets and is thus widely applicable in medical practice. The key idea of our method is that we leverage prediction-powered inferences and thereby essentially `shrink' the CIs so that we offer more precise uncertainty quantification as compared to naïve approaches. We further prove the unbiasedness of our method and the validity of our CIs. We confirm our theoretical results through various numerical experiments. Finally, we provide an extension of our method for constructing CIs from combinations of experimental and observational datasets.

因果推断置信区间医疗数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。