用自注意力与扩散模型处理部分断电下的多变量时间序列缺失数据。
Self-attention-based Diffusion Model for Time-series Imputation in Partial Blackout Scenarios
- 分两阶段利用自注意力和扩散过程建模特征与时间相关性。
- 在真实数据集上优于当前最佳方法,且支持不完整数据训练。
- 适合处理传感器断连、部分数据丢失等工业场景的补全任务。
多变量时间序列中的缺失值会损害机器学习性能并引入偏差,通常由传感器故障、停电或人为错误导致。现有方法主要针对随机缺失或完全断电场景。本文提出一种新的“部分断电”模式——连续时间段内部分特征缺失。为此设计了一种基于自注意力与扩散过程的两阶段插补方法,有效建模特征间与时间上的相关性。模型在训练中即可处理缺失数据,提升适应性,确保在不完整数据下仍具可靠插补能力与性能。在基准数据集及两个真实世界时间序列数据集上的实验表明,该方法在部分断电场景下超越当前最优水平,并展现出更优可扩展性。
原文摘要 · Abstract (English)
Missing values in multivariate time series data can harm machine learning performance and introduce bias. These gaps arise from sensor malfunctions, blackouts, and human error and are typically addressed by data imputation. Previous work has tackled the imputation of missing data in random, complete blackouts and forecasting scenarios. The current paper addresses a more general missing pattern, which we call "partial blackout," where a subset of features is missing for consecutive time steps. We introduce a two-stage imputation process using self-attention and diffusion processes to model feature and temporal correlations. Notably, our model effectively handles missing data during training, enhancing adaptability and ensuring reliable imputation and performance, even with incomplete datasets. Our experiments on benchmark and two real-world time series datasets demonstrate that our model outperforms the state-of-the-art in partial blackout scenarios and shows better scalability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。