arXiv:2505.14421stat.MLcs.LG2025-05

用系统辨识方法聚类向量自回归时间序列,避免依赖人工特征。

A system identification approach to clustering vector autoregressive time series

  • 基于混合自回归模型构建聚类算法,显式建模动态特性。
  • 提出k-LMVAR算法,在小噪声下计算效率高,仿真表现优异。
  • 适合需要自动聚类复杂时间序列的科研与工程场景。

基于底层动态特性的时间序列聚类因其在复杂系统建模中的作用而持续受到关注。现有方法多局限于标量时间序列,常将其视为白噪声,或依赖领域知识进行高质量特征构造,往往忽略自相关模式。本文提出一种系统辨识方法,通过显式考虑向量时间序列的自回归动态实现聚类。首先基于混合自回归模型推导出聚类算法,但存在显著计算瓶颈。随后提出其‘小噪声’极限版本——k-LMVAR(Limiting Mixture Vector AutoRegression),计算上可管理。开发了配套的BIC准则以确定簇数和模型阶数。该算法在对比模拟中表现优异,且计算可扩展性强。

原文摘要 · Abstract (English)

Clustering of time series based on their underlying dynamics is keeping attracting researchers due to its impacts on assisting complex system modelling. Most current time series clustering methods handle only scalar time series, treat them as white noise, or rely on domain knowledge for high-quality feature construction, where the autocorrelation pattern/feature is mostly ignored. Instead of relying on heuristic feature/metric construction, the system identification approach allows treating vector time series clustering by explicitly considering their underlying autoregressive dynamics. We first derive a clustering algorithm based on a mixture autoregressive model. Unfortunately it turns out to have significant computational problems. We then derive a `small-noise' limiting version of the algorithm, which we call k-LMVAR (Limiting Mixture Vector AutoRegression), that is computationally manageable. We develop an associated BIC criterion for choosing the number of clusters and model order. The algorithm performs very well in comparative simulations and also scales well computationally.

时间序列聚类系统辨识向量自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。