构建可复现的模拟数据集,评估精神科药物试验中合成对照组的准确性与可靠性。
Psych-ECA: A Reproducible Semi-Synthetic Benchmark for Synthetic Control Arms in Longitudinal Psychiatry
- 设计半合成数据生成器,模拟抑郁/焦虑/精神病症状轨迹及非随机随访。
- 轨迹建模方法误差仅2.3(PHQ-9 RMSE),Scribe模型兼具高准确率与校准性。
- 适合关注真实世界研究中因果推断与试验决策可靠性的研究人员。
外部与合成对照组正进入精神科药物研发,但缺乏能评估监管机构关心属性的基准:不仅需准确重建未治疗轨迹,还需不确定性校准、对精神健康记录中常见的‘病情越重随访越频繁’现象具有鲁棒性,以及在决策中保持低假阳性率。真实精神科试验数据(如STAR-D和注册队列)需认证访问且缺乏真实反事实数据,因此我们借鉴因果推断领域的半合成基准(IHDP、ACIC、PK-PD肿瘤生长模拟器),发布Psych-ECA——一个完全可复现的纵向症状轨迹生成器,涵盖抑郁症(PHQ-9)、焦虑症(HAM-A)和精神病(PANSS),包含已知反事实对照组、信息性随访和经验证的测量噪声。我们评估了八种估计器:前向携带、真实世界数据平均、最近邻匹配、线性混合模型、梯度提升、Scribe轨迹桥接法等。结果表明:第一,轨迹与灵活机器学习方法取得最佳反事实精度(约2.3 PHQ-9 RMSE),优于截面基线;第二,仅Scribe同时具备高准确率与校准性,其90%置信区间的实际覆盖率为93-96%,而梯度提升为87-88%,未校准的SDE模型仅为62-75%;第三,逆强度校正可减少信息性采样下的偏倚,而只有Scribe的校准区间能在信息性增强时维持名义假阳性率。所有代码、数据生成脚本及随机种子均已公开,支持完全可复现评估。
原文摘要 · Abstract (English)
External and synthetic control arms (ECAs) are entering psychiatric drug development, but the field lacks a benchmark that evaluates the properties regulators care about: not only how accurately a method reconstructs untreated trajectories, but whether its uncertainty is calibrated, whether it is robust to the informative observation times common in mental-health records (sicker patients are seen more often), and what false-positive rate it induces in go/no-go trial decisions. Real psychiatric trial data (e.g. STAR-D and registry cohorts) require credentialed access and lack ground-truth counterfactuals, so, following established semi-synthetic benchmarks in causal inference (IHDP, ACIC, and the PK-PD tumor-growth simulator), we release Psych-ECA, a fully reproducible generator of longitudinal symptom trajectories for depression (PHQ-9), anxiety (HAM-A), and psychosis (PANSS) with known counterfactual control arms, informative visits, and validated-scale measurement noise. We benchmark eight estimators spanning carry-forward, pooled real-world-data averages, nearest-neighbour matching, linear mixed models, gradient boosting, and the Scribe trajectory-bridge method. Three findings emerge. First, trajectory and flexible machine learning methods achieve the best counterfactual accuracy (about 2.3 PHQ-9 RMSE), outperforming cross-sectional baselines. Second, only Scribe is both accurate and calibrated, achieving 93-96% empirical coverage of nominal 90% prediction intervals, compared with 87-88% for gradient boosting and 62-75% for uncalibrated SDE models. Third, inverse-intensity correction reduces bias under informative sampling, while Scribe's calibrated intervals are the only trajectory method that maintains nominal false-positive rates as informativeness increases. We release all code, data-generation scripts, and random seeds to enable fully reproducible evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。