arXiv:2505.11014stat.MEcs.LG2025-05NeurIPS

融合不同评估指标的研究需谨慎,否则可能引入偏倚。

A Cautionary Tale on Integrating Studies with Disparate Outcome Measures for Causal Inference

  • 通过三类假设关联不同量表,分析数据整合的可行性
  • 仅在最强假设下能提升渐近效率,否则易产生偏倚
  • 小样本时或有增效,但大样本下优势消失

数据整合方法日益用于提升研究效率与普适性,但其关键假设是各数据集结果度量一致——这在实践中常不成立。以阿片类药物使用障碍(OUD)研究为例,XBOT试验与POAT研究均评估药物对戒断症状严重程度的影响(非主要终点),但分别采用主观阿片戒断量表(SOW)与临床阿片戒断量表(COW)。本文分析了结果量表不一致且无交叉记录的现实场景。提出三组不同程度的假设来关联两种量表,理论与实证结果显示:仅在最强假设下可实现渐近效率提升,否则会引入偏倚;较弱假设虽可在小样本中带来有限效率增益,但随样本量增大而减弱。通过整合XBOT与POAT数据估计两种药物对戒断症状的相对效果,系统变化假设条件,揭示了效率增益与偏倚风险之间的权衡。研究强调在整合不同结果量表的数据时,必须审慎选择假设,为现代数据融合提供重要指引。

原文摘要 · Abstract (English)

Data integration approaches are increasingly used to enhance the efficiency and generalizability of studies. However, a key limitation of these methods is the assumption that outcome measures are identical across datasets -- an assumption that often does not hold in practice. Consider the following opioid use disorder (OUD) studies: the XBOT trial and the POAT study, both evaluating the effect of medications for OUD on withdrawal symptom severity (not the primary outcome of either trial). While XBOT measures withdrawal severity using the subjective opiate withdrawal scale, POAT uses the clinical opiate withdrawal scale. We analyze this realistic yet challenging setting where outcome measures differ across studies and where neither study records both types of outcomes. Our paper studies whether and when integrating studies with disparate outcome measures leads to efficiency gains. We introduce three sets of assumptions -- with varying degrees of strength -- linking both outcome measures. Our theoretical and empirical results highlight a cautionary tale: integration can improve asymptotic efficiency only under the strongest assumption linking the outcomes. However, misspecification of this assumption leads to bias. In contrast, a milder assumption may yield finite-sample efficiency gains, yet these benefits diminish as sample size increases. We illustrate these trade-offs via a case study integrating the XBOT and POAT datasets to estimate the comparative effect of two medications for opioid use disorder on withdrawal symptoms. By systematically varying the assumptions linking the SOW and COW scales, we show potential efficiency gains and the risks of bias. Our findings emphasize the need for careful assumption selection when fusing datasets with differing outcome measures, offering guidance for researchers navigating this common challenge in modern data integration.

因果推断数据融合偏倚控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。