用数字孪生生成工业用电数据集,提升复杂设备能耗分解精度
Industrial Energy Disaggregation with Digital Twin-generated Dataset and Efficient Data Augmentation
- 通过数字孪生模拟生成三类工业设施的合成数据集
- 新方法使模型在真实场景下误差降至0.093,优于传统方法
- 适合工业能源管理、智能电网研究者使用
工业非侵入式负荷监测受限于高质量数据稀缺和能耗模式复杂多变。为解决数据不足与隐私问题,我们提出面向能耗分解的合成工业数据集(SIDED),基于数字孪生仿真生成,涵盖三种工业设施类型及三个地理区域,包含多样化的设备行为、天气条件与负荷曲线。同时提出设备调制数据增强(AMDA)方法,通过按设备相对影响智能调节功率贡献,实现高效计算的数据增强。实验表明,采用AMDA增强数据训练的模型,在跨样本场景中能耗分解误差达0.093,显著优于无增强(0.451)和随机增强(0.290)。数据分布分析证实,AMDA有效对齐训练与测试数据分布,提升模型泛化能力。
原文摘要 · Abstract (English)
Industrial Non-Intrusive Load Monitoring (NILM) is limited by the scarcity of high-quality datasets and the complex variability of industrial energy consumption patterns. To address data scarcity and privacy issues, we introduce the Synthetic Industrial Dataset for Energy Disaggregation (SIDED), an open-source dataset generated using Digital Twin simulations. SIDED includes three types of industrial facilities across three different geographic locations, capturing diverse appliance behaviors, weather conditions, and load profiles. We also propose the Appliance-Modulated Data Augmentation (AMDA) method, a computationally efficient technique that enhances NILM model generalization by intelligently scaling appliance power contributions based on their relative impact. We show in experiments that NILM models trained with AMDA-augmented data significantly improve the disaggregation of energy consumption of complex industrial appliances like combined heat and power systems. Specifically, in our out-of-sample scenarios, models trained with AMDA achieved a Normalized Disaggregation Error of 0.093, outperforming models trained without data augmentation (0.451) and those trained with random data augmentation (0.290). Data distribution analyses confirm that AMDA effectively aligns training and test data distributions, enhancing model generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。