arXiv:2602.02371cs.LGstat.ML2026-02

用近邻搜索分析长期疾病数据,精准估算治疗效果差异。

C-kNN-LSH: A Nearest-Neighbor Algorithm for Sequential Counterfactual Inference

  • 基于局部敏感哈希找相似病程患者,实现动态病情下的因果推断。
  • 在13,511人的真实长新冠数据中,准确捕捉恢复差异并评估干预政策价值。
  • 对不规则采样和恢复轨迹变化有强鲁棒性,适合临床决策支持场景。

从纵向轨迹中估计因果效应是理解复杂疾病进展与优化临床决策的核心,如共病及长新冠康复。我们提出C-kNN-LSH,一种针对序列因果推断的近邻框架,用于处理高维、混杂严重的场景。通过局部敏感哈希(LSH),高效识别具有相似协变量历史的“临床孪生”个体,实现对随疾病状态演变的条件治疗效应的局部估计。为缓解不规则采样和患者恢复轨迹变化带来的偏差,引入邻域估计器与双重稳健校正。理论分析证明该估计器一致且对扰动误差二阶稳健。在包含13,511名参与者的现实长新冠队列上评估,相比现有基线,C-kNN-LSH在捕捉恢复异质性和估计政策价值方面表现更优。

原文摘要 · Abstract (English)

Estimating causal effects from longitudinal trajectories is central to understanding the progression of complex conditions and optimizing clinical decision-making, such as comorbidities and long COVID recovery. We introduce \emph{C-kNN--LSH}, a nearest-neighbor framework for sequential causal inference designed to handle such high-dimensional, confounded situations. By utilizing locality-sensitive hashing, we efficiently identify ``clinical twins'' with similar covariate histories, enabling local estimation of conditional treatment effects across evolving disease states. To mitigate bias from irregular sampling and shifting patient recovery profiles, we integrate neighborhood estimator with a doubly-robust correction. Theoretical analysis guarantees our estimator is consistent and second-order robust to nuisance error. Evaluated on a real-world Long COVID cohort with 13,511 participants, \emph{C-kNN-LSH} demonstrates superior performance in capturing recovery heterogeneity and estimating policy values compared to existing baselines.

因果推断长新冠近邻算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。