区分缺失机制,提升时间序列填补准确率
Causal View of Time Series Imputation: Some Identification Results on Missing Mechanism
- 根据缺失机制差异设计针对性填补方法
- 在多种数据集上超越现有技术,有效应对真实场景
- 适合处理医疗、物联网等含复杂缺失的数据
时间序列填补是医疗、物联网等领域的重要挑战。现有方法通常忽略缺失机制差异,统一建模导致结果偏差。本文提出基于不同缺失机制(DMM)的填补框架,分析不同机制下的生成过程,利用变分推断与基于归一化流的神经架构建模潜在状态与缺失原因,并在非线性独立成分分析框架下建立可识别性结果,证明潜在变量可被识别。实验表明,该方法在多种数据集和缺失机制下均优于现有技术,具备实际应用价值。
原文摘要 · Abstract (English)
Time series imputation is one of the most challenge problems and has broad applications in various fields like health care and the Internet of Things. Existing methods mainly aim to model the temporally latent dependencies and the generation process from the observed time series data. In real-world scenarios, different types of missing mechanisms, like MAR (Missing At Random), and MNAR (Missing Not At Random) can occur in time series data. However, existing methods often overlook the difference among the aforementioned missing mechanisms and use a single model for time series imputation, which can easily lead to misleading results due to mechanism mismatching. In this paper, we propose a framework for time series imputation problem by exploring Different Missing Mechanisms (DMM in short) and tailoring solutions accordingly. Specifically, we first analyze the data generation processes with temporal latent states and missing cause variables for different mechanisms. Sequentially, we model these generation processes via variational inference and estimate prior distributions of latent variables via normalizing flow-based neural architecture. Furthermore, we establish identifiability results under the nonlinear independent component analysis framework to show that latent variables are identifiable. Experimental results show that our method surpasses existing time series imputation techniques across various datasets with different missing mechanisms, demonstrating its effectiveness in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。