arXiv:2412.11164cs.LGstat.AP2024-12被引 4

MICE-RF在噪声时间序列缺失值填补中表现优于深度学习模型,兼具去噪效果。

Missing data imputation for noisy time-series data and applications in healthcare

  • 采用MICE-RF等方法对比填补不同缺失率下的医疗时间序列数据
  • 在10%-80%缺失率下,MICE-RF的MAE更低,分类性能更优
  • 填补过程本身具有去噪作用,适合临床数据预处理

医疗时间序列数据对监测患者状态至关重要,但常因传感器故障或数据中断导致噪声和缺失。填补缺失值(imputation)是常见应对策略。本研究比较了多重插补随机森林(MICE-RF)与先进深度学习方法(SAITS、BRITS、Transformer)在噪声缺失时间序列数据上的表现,评估指标包括MAE、F1-score、AUC和MCC,覆盖10%至80%的缺失率。结果表明,相比深度学习方法,MICE-RF在缺失值填补上更具优势,且填补后数据的分类性能提升,说明插补过程具有去噪效应。因此,在存在缺失值的时间序列上使用插补算法,可同时实现数据修复与降噪。

原文摘要 · Abstract (English)

Healthcare time series data is vital for monitoring patient activity but often contains noise and missing values due to various reasons such as sensor errors or data interruptions. Imputation, i.e., filling in the missing values, is a common way to deal with this issue. In this study, we compare imputation methods, including Multiple Imputation with Random Forest (MICE-RF) and advanced deep learning approaches (SAITS, BRITS, Transformer) for noisy, missing time series data in terms of MAE, F1-score, AUC, and MCC, across missing data rates (10 % - 80 %). Our results show that MICE-RF can effectively impute missing data compared to deep learning methods and the improvement in classification of data imputed indicates that imputation can have denoising effects. Therefore, using an imputation algorithm on time series with missing data can, at the same time, offer denoising effects.

时间序列缺失值填补医疗数据去噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。