系统研究Transformer在家庭用电监测中的超参影响,找到最优配置提升性能。
Towards a Deeper Understanding of Transformer for Residential Non-intrusive Load Monitoring
- 测试隐藏维度、注意力层、头数和丢弃率对模型的影响
- 发现最佳配置下模型超越现有方法,显著提升识别准确率
- 为电力负荷分析中的Transformer设计提供实证指导
近年来,Transformer模型在非侵入式负载监测(NILM)任务中表现出色。尽管如此,现有研究尚未充分探讨各类超参数对模型性能的影响,而这对于构建高性能Transformer模型至关重要。本文通过一系列全面实验,分析了注意力层隐藏维度数量、注意力层数量、注意力头数以及丢弃率对住宅场景下NILM性能的影响。此外,还研究了BERT风格训练中掩码比例的作用,深入评估其在NILM任务中的影响。基于实验结果,确定了最优超参数组合并据此训练出的Transformer模型,在性能上优于现有模型。这些发现为优化Transformer架构提供了宝贵洞见与实用指南,旨在提升其在NILM应用中的有效性与效率。本工作有望成为未来更鲁棒、更强大变压器模型研发的基础。
原文摘要 · Abstract (English)
Transformer models have demonstrated impressive performance in Non-Intrusive Load Monitoring (NILM) applications in recent years. Despite their success, existing studies have not thoroughly examined the impact of various hyper-parameters on model performance, which is crucial for advancing high-performing transformer models. In this work, a comprehensive series of experiments have been conducted to analyze the influence of these hyper-parameters in the context of residential NILM. This study delves into the effects of the number of hidden dimensions in the attention layer, the number of attention layers, the number of attention heads, and the dropout ratio on transformer performance. Furthermore, the role of the masking ratio has explored in BERT-style transformer training, providing a detailed investigation into its impact on NILM tasks. Based on these experiments, the optimal hyper-parameters have been selected and used them to train a transformer model, which surpasses the performance of existing models. The experimental findings offer valuable insights and guidelines for optimizing transformer architectures, aiming to enhance their effectiveness and efficiency in NILM applications. It is expected that this work will serve as a foundation for future research and development of more robust and capable transformer models for NILM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。