处理不规则时间序列数据的因果推断新方法,能有效应对高维混杂和信息性测量。
Doubly Robust Functional Representation Learning for Longitudinal Causal Inference with Irregular Histories

- 将不规则历史数据映射为针对目标估计量优化的状态表示
- 在高维混杂、弱支持、重尾伪结果下提升估计精度
- 适合重症监护等存在不规则监测数据的医学因果分析
纵向因果研究常记录不规则的功能性历史数据:如实验室值、生理信号、传感器流、影像摘要等,在不等且具信息量的时间点采集。标准双重稳健估计器通常依赖标量汇总,而序列学习优化预测损失,未必稳定高效影响函数。本文提出双重稳健功能表示学习(DR-FRL),一种交叉拟合流程,将不规则历史转化为观测历史制度下的目标估计量相关状态。功能与时间编码器将点云和历史映射为状态;干扰头估计结果、治疗及删失函数;基于高效影响函数的目标验证、校准、重叠、尾部及消融诊断评估状态是否支持估计方程。若所选状态保留了高效影响函数所需的干扰信息,则表示误差进入与普通干扰误差相同的二阶乘积余项,均值估计器在明确速率、重叠、校准与稳定性条件下渐近线性。Catoni聚合单独作为有界影响点估计,非沃尔德推断替代。模拟显示在高维混杂、信息性测量、支持弱或伪结果重尾时性能提升。对VitalDB的审计表明,DR-FRL可利用不规则实验室点云,得出有用负结果:对于此重症监护室转归终点,标量实验室汇总已包含大量终点相关信息。
原文摘要 · Abstract (English)
Longitudinal causal studies often record histories as irregular functional fragments: laboratory values, physiologic signals, sensor streams, and image-derived summaries measured at unequal and informative times. Standard doubly robust estimators usually require scalar summaries, whereas sequence learners optimize prediction losses that need not stabilize the efficient influence function. We propose Doubly Robust Functional Representation Learning (DR-FRL), a cross-fitted workflow that turns irregular histories into estimand-targeted states for observed-history regimes. Functional and temporal encoders map point clouds and prior histories into states; nuisance heads estimate outcome, treatment, and censoring functions; and EIF-targeted validation, calibration, overlap, tail, and ablation diagnostics assess whether the state supports the estimating equation. If the selected state preserves the nuisance information needed by the EIF, representation error enters the same second-order product remainder as ordinary nuisance error, and the mean estimator is asymptotically linear under explicit rate, overlap, calibration, and stability conditions. Catoni aggregation is treated separately as a bounded-influence point estimator, not a replacement for Wald inference. Simulations show gains when functional confounding is high-dimensional, measurement is informative, support is weak, or pseudo-outcomes are heavy-tailed. A VitalDB audit shows that DR-FRL can use irregular laboratory point clouds and deliver a useful negative finding: for this ICU-disposition endpoint, scalar laboratory summaries already carry much endpoint-relevant information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。