arXiv:2510.24217cs.LGcs.AI2025-10被引 2

对比15种插补方法,找出重症监护数据最有效的填补策略。

Closing Gaps: An Imputation Analysis of ICU Vital Signs

  • 构建可扩展的基准测试框架,涵盖15种插补与4种伪造缺失方法。
  • 实证发现零值插补显著降低预测性能,需谨慎使用。
  • 为临床预测模型研究者提供可复用的插补方法选择依据。

随着重症监护室(ICU)数据日益丰富,利用机器学习开发临床预测模型以优化医疗流程的兴趣不断增长。然而,数据质量不足仍是制约临床预测应用的关键问题。许多生命体征测量(如心率)存在大量缺失片段,形成数据空白,可能严重影响预测效果。尽管已有多种时间序列插补技术被提出,但尚缺乏对代表性方法的系统性比较,难以确定最佳实践。现实中,仍存在使用零值插补等会降低准确性的随意插补方式。本文通过对比主流插补方法,指导研究者通过选择更优插补策略提升临床预测模型性能。我们构建了一个可扩展、可复用的基准测试平台,包含当前15种插补方法和4种模拟缺失方法,适用于主要ICU数据集。旨在提供一个可比基础,推动更多模型走向临床实践。

原文摘要 · Abstract (English)

As more Intensive Care Unit (ICU) data becomes available, the interest in developing clinical prediction models to improve healthcare protocols increases. However, the lack of data quality still hinders clinical prediction using Machine Learning (ML). Many vital sign measurements, such as heart rate, contain sizeable missing segments, leaving gaps in the data that could negatively impact prediction performance. Previous works have introduced numerous time-series imputation techniques. Nevertheless, more comprehensive work is needed to compare a representative set of methods for imputing ICU vital signs and determine the best practice. In reality, ad-hoc imputation techniques that could decrease prediction accuracy, like zero imputation, are still used. In this work, we compare established imputation techniques to guide researchers in improving the performance of clinical prediction models by selecting the most accurate imputation technique. We introduce an extensible and reusable benchmark with currently 15 imputation and 4 amputation methods, created for benchmarking on major ICU datasets. We hope to provide a comparative basis and facilitate further ML development to bring more models into clinical practice.

ICU数据时间序列插补方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。