用概率嵌入聚类光伏系统,提升数据缺失时的分析可靠性。
Clustering Rooftop PV Systems via Probabilistic Embeddings
- 将光伏系统发电特征与不确定性编码为概率分布进行聚类
- 在多时间跨度住宅数据集上实现更鲁棒的聚类表现
- 适合需要处理缺失数据的能源系统分析人员
随着屋顶光伏安装数量增加,聚合商和系统运营商需监控和分析这些分布式时序数据,面临高维、空间分散且含缺失值的挑战。本文提出基于概率实体嵌入的聚类框架:将每个光伏系统的发电模式与不确定性编码为概率分布,通过统计距离与层次聚类进行分组。应用于多年住宅光伏数据集后,生成简洁且具备不确定性的聚类画像,其代表性与鲁棒性优于基于物理模型的基线方法,并支持可靠的缺失值填补。系统性超参数研究进一步为平衡模型性能与鲁棒性提供实用指导。
原文摘要 · Abstract (English)
As the number of rooftop photovoltaic (PV) installations increases, aggregators and system operators are required to monitor and analyze these systems, raising the challenge of integration and management of large, spatially distributed time-series data that are both high-dimensional and affected by missing values. In this work, a probabilistic entity embedding-based clustering framework is proposed to address these problems. This method encodes each PV system's characteristic power generation patterns and uncertainty as a probability distribution, then groups systems by their statistical distances and agglomerative clustering. Applied to a multi-year residential PV dataset, it produces concise, uncertainty-aware cluster profiles that outperform a physics-based baseline in representativeness and robustness, and support reliable missing-value imputation. A systematic hyperparameter study further offers practical guidance for balancing model performance and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。