arXiv:2506.20023cs.LGcs.DB2025-06被引 3

针对真实场景的复杂缺损数据,提出动态填补框架DIM-SUM,提升预测精度与效率。

DIM-SUM: Dynamic IMputation for Smart Utility Management

  • 基于模式聚类与自适应掩码,模拟真实缺失模式训练模型
  • 在超20亿条加州水电气数据上,用更少数据达更高精度
  • 相比大模型快2倍,适合工业级实时监控系统

时间序列填补模型传统上使用人工掩码的完整数据集进行训练。然而在真实基础设施监控中,数据常出现大量缺失且模式复杂多变。我们提出DIM-SUM,一种用于训练鲁棒填补模型的预处理框架,弥合了人工掩码数据与真实缺失模式之间的差距。DIM-SUM结合模式聚类与自适应掩码策略,并具备理论学习保证,可处理实际数据中观察到的多样化缺失模式。在来自加州水务区、电力数据集及基准测试的超20亿条读数上进行的广泛实验表明,与传统方法相比,DIM-SUM在更低的处理时间和更少训练数据下达到相近甚至更高的准确率;相较于大型预训练模型,其平均精度高出2倍,推理时间显著更短。

原文摘要 · Abstract (English)

Time series imputation models have traditionally been developed using complete datasets with artificial masking patterns to simulate missing values. However, in real-world infrastructure monitoring, practitioners often encounter datasets where large amounts of data are missing and follow complex, heterogeneous patterns. We introduce DIM-SUM, a preprocessing framework for training robust imputation models that bridges the gap between artificially masked training data and real missing patterns. DIM-SUM combines pattern clustering and adaptive masking strategies with theoretical learning guarantees to handle diverse missing patterns actually observed in the data. Through extensive experiments on over 2 billion readings from California water districts, electricity datasets, and benchmarks, we demonstrate that DIM-SUM outperforms traditional methods by reaching similar accuracy with lower processing time and significantly less training data. When compared against a large pre-trained model, DIM-SUM averages 2x higher accuracy with significantly less inference time.

时间序列填补智能电网数据缺失

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。