arXiv:2505.02173stat.MLcs.LG2025-05被引 2

提出新相似性度量,更好区分电力用户用电模式差异。

Ranked differences Pearson correlation dissimilarity with an application to electricity users time series clustering

  • 结合最大差值加权平均与皮尔逊相关性,设计新距离度量。
  • 在含季节、趋势、峰谷的复杂数据中聚类效果更优。
  • 适用于电力消费等具有多特征时间序列的用户分群。

时间序列聚类是一种无监督学习方法,用于将时间序列数据划分为行为相似的组,广泛应用于医疗、金融、能源和气候科学等领域。过去四十年间已有多种聚类方法被提出,多数聚焦于欧氏距离或关联性差异的度量。本文提出一种新的相异度量——排名皮尔逊相关相异度(RDPC),该方法结合特定比例的最大逐元素差值加权平均与经典的皮尔逊相关相异度。将其嵌入层次聚类框架中进行评估,并与现有算法对比。结果表明,在包含不同季节模式、趋势及峰值的复杂情形下,RDPC表现更优。最后,我们将该方法应用于泰国电力消费时间序列数据集中的随机用户样本,成功将其聚类为七个具有独特特征的群体。

原文摘要 · Abstract (English)

Time series clustering is an unsupervised learning method for classifying time series data into groups with similar behavior. It is used in applications such as healthcare, finance, economics, energy, and climate science. Several time series clustering methods have been introduced and used for over four decades. Most of them focus on measuring either Euclidean distances or association dissimilarities between time series. In this work, we propose a new dissimilarity measure called ranked Pearson correlation dissimilarity (RDPC), which combines a weighted average of a specified fraction of the largest element-wise differences with the well-known Pearson correlation dissimilarity. It is incorporated into hierarchical clustering. The performance is evaluated and compared with existing clustering algorithms. The results show that the RDPC algorithm outperforms others in complicated cases involving different seasonal patterns, trends, and peaks. Finally, we demonstrate our method by clustering a random sample of customers from a Thai electricity consumption time series dataset into seven groups with unique characteristics.

时间序列聚类电力分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。