用隐式神经表示填补缺失医疗数据,提升疾病诊断效果
ImputeINR: Time Series Imputation via Implicit Neural Representations for Disease Diagnosis with Missing Data
- 用隐式神经网络建模时间序列连续函数,不依赖采样频率
- 在缺失率高达80%时仍能生成高精度填补结果
- 适用于医疗数据缺失严重场景,提升下游诊断模型表现
医疗数据常存在大量缺失值,需有效的时间序列填补方法以支持疾病诊断。现有方法多针对离散数据点,难以建模稀疏数据,尤其在缺失比例高时性能显著下降。本文提出ImputeINR方法,利用隐式神经表示(INR)学习时间序列的连续函数,其连续表示不依赖采样频率且具有无限采样能力,可在极稀疏观测下生成精细填补。在八个数据集上、五种缺失比例条件下进行的大量实验表明,ImputeINR在高缺失率下均表现更优。进一步验证显示,使用ImputeINR填补医疗数据可显著提升下游疾病诊断任务性能。代码已公开。
原文摘要 · Abstract (English)
Healthcare data frequently contain a substantial proportion of missing values, necessitating effective time series imputation to support downstream disease diagnosis tasks. However, existing imputation methods focus on discrete data points and are unable to effectively model sparse data, resulting in particularly poor performance for imputing substantial missing values. In this paper, we propose a novel approach, ImputeINR, for time series imputation by employing implicit neural representations (INR) to learn continuous functions for time series. ImputeINR leverages the merits of INR in that the continuous functions are not coupled to sampling frequency and have infinite sampling frequency, allowing ImputeINR to generate fine-grained imputations even on extremely sparse observed values. Extensive experiments conducted on eight datasets with five ratios of masked values show the superior imputation performance of ImputeINR, especially for high missing ratios in time series data. Furthermore, we validate that applying ImputeINR to impute missing values in healthcare data enhances the performance of downstream disease diagnosis tasks. Codes are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。