提出高效特征选择方法CSFS,提升风电光伏预测精度与效率
Improving Wind and Solar Power Prediction with Efficient Wrapper-based Feature Selection: An Empirical Study
- 基于聚类的包装器方法,自动筛选关键输入特征
- 相比传统方法,计算成本降低21%且预测性能相当
- 适合能源预测研究者与实际应用开发者参考
随着全球能源需求上升及对气候变化影响的关注增加,可再生能源在能源结构中的占比持续增长。由于依赖环境条件,可再生能源输出难以像传统发电那样稳定可控,因此准确预测当前和未来发电量至关重要。本文通过两项系统性文献综述,分析了风力涡轮机功率曲线建模和光伏发电预测两大实际任务。研究发现,尽管监测与环境变量众多,但可用的特征选择方法仍有限且缺乏系统性。为此,我们提出一种新型、模型无关、基于聚类的包装器方法——簇基序列特征选择(CSFS),实现可再生能源预测流程中自动、高效、可靠的特征选择。为支持可复现性与再利用,我们在GitHub上开源了该方法。在两个应用场景中,我们对所提方法进行了实证评估,并与封装式序列特征选择(SFS)、过滤式方法及随机森林内置重要性等主流方法进行比较。结果表明,封装式方法整体表现更优;而CSFS在达到与SFS相当的预测性能的同时,平均降低了21%的计算开销。
原文摘要 · Abstract (English)
With rising global energy demand and growing awareness of climate change and its impacts, the share of renewable energies in the global energy mix continues to grow. Unlike conventional power generation, the output of renewable energy sources cannot be controlled as consistently due to their dependence on environmental conditions. Therefore, reliable prediction of current and future energy production is essential. In this paper, we report findings from two structured literature reviews on real-world renewable energy prediction tasks: wind turbine power curve modeling and photovoltaic power prediction. For the former, we conducted a comprehensive literature review ourselves, while for the latter, we synthesize the key findings regarding frequently selected input features based on an existing survey. Across both domains, our analysis reveals that despite the large number of available monitoring and environmental variables, only limited or unsystematic methods for feature selection exist. To address this gap, we propose Cluster-based Sequential Feature Selection (CSFS), a novel, model-agnostic, clustering-based wrapper method for automatic, efficient, and reliable feature selection in renewable energy prediction pipelines. To support reproducibility and reuse, we provide an open-source implementation of CSFS on GitHub. We empirically evaluate the proposed approach on both use cases and compare it with established feature selection techniques such as wrapper-based sequential feature selection (SFS), filter-based methods, and Random Forest's embedded feature importance. The results show that the wrapper-based methods overall provide better-performing selections of features. CSFS achieves a predictive performance comparable to SFS while reducing computational cost by an average of 21%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。