arXiv:2607.05613cs.LG2026-07KDD

提出可信赖临床数据补全方法,确保高风险场景下结果可靠

SafeImpute: Reliable Clinical Data Imputation via Conformal Selection

论文配图:SafeImpute: Reliable Clinical Data Imputation via Conformal Selection
图 1 · 摘自论文原文
  • 构建患者时序与跨患者相似性图,用双关系GNN融合补全
  • 在梅奥诊所、MIMIC-III/IV数据上准确率领先且错误可控
  • 适合医疗决策等高风险应用,提供可解释的可靠性保障

临床诊疗依赖关键检验指标,但实际就诊中检查稀疏且不规律,导致数据缺失普遍。现有补全方法虽提升平均精度,却难以判断哪些结果足够可靠用于高风险下游任务。本文提出SafeImpute,一种面向不规则稀疏临床纵向记录的可靠补全框架。该方法构建事件图,捕捉患者内部时序轨迹与跨患者临床相似性,通过双关系图神经网络与自适应融合学习补全,并以掩码重建辅助目标正则化。为实现可靠性保障,将代理风险得分转化为置信区间p值,采用Benjamini–Hochberg程序控制用户指定容忍度下的不可接受误差的假发现率(FDR)。在梅奥诊所及公开的MIMIC-III、MIMIC-IV数据集上的实验表明,SafeImpute在标准补全评估和受控释放评估中均优于多种基线,兼具高精度与可靠的错误控制能力。

原文摘要 · Abstract (English)

Clinical care often relies on key laboratory indicators, yet real-world patient visits are sparse and tests are ordered irregularly, leading to pervasive missingness. While many imputation methods improve average accuracy, they provide limited guidance on which imputed values are reliable enough for high-stakes downstream use. In this work, we study reliable clinical imputation, aiming to produce accurate imputations while selectively releasing the reliable results, with statistical control over clinically unacceptable errors. To achieve this goal, we propose SafeImpute, a reliable imputation framework for irregular and sparse clinical longitudinal records. SafeImpute constructs an event graph that captures both intra-patient temporal trajectories and inter-patient clinical similarity, and learns imputations with a two-relation GNN and adaptive fusion, regularized by an auxiliary masked reconstruction objective. For reliability guarantees, SafeImpute converts a proxy risk score into conformal p-values and applies the Benjamini--Hochberg procedure to control the false discovery rate (FDR) of unacceptable errors among released imputations at a user-specified tolerance. Experiments on our Mayo Clinic data, the public MIMIC-III and MIMIC-IV datasets show that SafeImpute achieves strong imputation accuracy while providing reliable error control, outperforming diverse baselines in both standard imputation evaluation and FDR-controlled selective-release evaluation.

临床数据补全可靠性图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。