对比多种模型填补智能电表数据空缺,发现时序基础模型更准但更耗算力。
Bridging Smart Meter Gaps: A Benchmark of Statistical, Machine Learning and Time Series Foundation Models for Data Imputation
- 用时序基础模型和大语言模型填补智能电表数据缺口
- 时序模型在特定场景下准确率显著提升,最高达95.3%
- 适合需高精度补全的电网分析与预测任务
智能电网中的时间序列数据常因传感器故障、传输错误或中断而出现缺失值。电表数据的空白会扭曲用电分析结果并影响可靠预测,导致技术和经济效率下降。随着智能电表数据量与复杂性增加,传统方法难以应对非线性与非平稳特征。本文评估了两种通用大语言模型和五种时序基础模型在电表数据补全中的表现,对比了传统机器学习与统计模型。我们在匿名公开数据集上人为制造30分钟至一天的间隔缺失,测试模型推理能力。结果显示,具备上下文理解与模式识别能力的时序基础模型在部分情况下可显著提升补全精度。然而,计算成本与性能提升之间的权衡仍是关键考量。
原文摘要 · Abstract (English)
The integrity of time series data in smart grids is often compromised by missing values due to sensor failures, transmission errors, or disruptions. Gaps in smart meter data can bias consumption analyses and hinder reliable predictions, causing technical and economic inefficiencies. As smart meter data grows in volume and complexity, conventional techniques struggle with its nonlinear and nonstationary patterns. In this context, Generative Artificial Intelligence offers promising solutions that may outperform traditional statistical methods. In this paper, we evaluate two general-purpose Large Language Models and five Time Series Foundation Models for smart meter data imputation, comparing them with conventional Machine Learning and statistical models. We introduce artificial gaps (30 minutes to one day) into an anonymized public dataset to test inference capabilities. Results show that Time Series Foundation Models, with their contextual understanding and pattern recognition, could significantly enhance imputation accuracy in certain cases. However, the trade-off between computational cost and performance gains remains a critical consideration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。