arXiv:2605.05125cs.LGcs.AI2026-05

解决医疗数据缺失与时间依赖性问题,提升真实世界疗效评估准确性

Joint Treatment Effect Estimation from Incomplete Healthcare Data: Temporal Causal Normalizing Flows with LLM-driven Evolutionary MNAR Imputation

论文配图:Joint Treatment Effect Estimation from Incomplete Healthcare Data: Temporal Causal Normalizing Flows with LLM-driven Evolutionary MNAR Imputation
图 1 · 摘自论文原文
  • 用带因果图约束的可逆流模型建模患者随时间演变的健康轨迹
  • 在高达80%缺失数据下仍保持治疗效应估计准确,优于传统方法
  • 结合大模型自动补全缺失数据,适合真实医疗数据场景研究者使用

目标试验模拟(TTE)可在随机对照试验不可行时,利用观察数据回答因果问题。然而现有方法常将因果推断、缺失值处理和时间结构分开处理,在电子健康记录(EHR)中表现受限,其中时变混杂和缺失不随机(MNAR)生物标志物可达50%–80%。本文提出两阶段流程:首先,基于长短期记忆(LSTM)编码病史并受有向无环图(DAG)约束的可逆流模型(CausalFlow-T),实现精确反事实推断,避免变分推断近似误差,并通过显式因果结构分离混杂因素;在四个合成及一个半合成基准测试中验证,DAG约束与精确推断各解决不同失效模式,互不可替代。其次,针对输入需完整的要求,引入基于大语言模型(LLM)的演化式插补器,生成可执行的插补操作而非单一数值,采用三种LLM后端进行评估。在30%–80% MNAR缺失率下,该插补器在生物标志物与因果指标综合排名中表现最佳,点对点精度与时间外推能力突出,同时维持平均治疗效应(ATE)恢复稳定,当统计基线性能下降时仍具优势。在瑞士初级保健机构的2型糖尿病成人患者真实数据中,该流程估计出接受GLP-1受体激动剂或SGLT-2抑制剂患者的按方案减重差异为-0.98公斤[95% CI -1.01, -0.96],支持药物组更优,且结果来自高度不完整的现实数据。

原文摘要 · Abstract (English)

Target trial emulation (TTE) enables causal questions to be studied with observational data when randomized controlled trials (RCTs) are infeasible. Yet treatment-effect methods often address causal estimation, missingness, and temporal structure separately, limiting their robustness in electronic health records (EHRs), where time-varying confounding and missing-not-at-random (MNAR) biomarkers can reach 50%--80%. We propose a two-stage pipeline for treatment effect estimation from incomplete longitudinal EHRs. First, CausalFlow-T, a directed acyclic graph (DAG)-constrained normalizing flow with long short-term memory (LSTM)-encoded patient history, performs exact invertible counterfactual inference, avoiding approximation errors from variational inference and separating confounding through explicit causal structure. Ablations on four synthetic and one semi-synthetic benchmark with known counterfactuals show that DAG constraints and exact inference address distinct failure modes: neither compensates for the other. Second, because CausalFlow-T requires completed inputs, we introduce an LLM-driven evolutionary imputer that proposes executable imputation operators rather than individual entries, and evaluate it with three large language model (LLM) backends, including two open-source models. Across 30%--80% MNAR missingness, this imputer achieves the best pooled rank over biomarker and causal metrics, leading in point-wise accuracy and temporal extrapolation while preserving average treatment effect (ATE) recovery as statistical baselines degrade. On Swiss primary-care EHRs from adults with type 2 diabetes initiating a GLP-1 receptor agonist or SGLT-2 inhibitor, the pipeline estimates a per-protocol weight-loss difference of -0.98 kg [95% CI -1.01, -0.96] favoring GLP-1 receptor agonists, consistent with randomized evidence and obtained from realistically incomplete real-world EHRs.

因果推断医疗数据缺失值处理时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。