提出首个系统评估时间序列相关结构发现能力的基准
CSTS: A Benchmark for the Discovery of Correlation Structures in Time Series Clustering
- 构建可调控相关结构的合成数据集,分离算法与评估方法影响
- 验证降采样导致中等程度结构失真,分布偏移影响小
- 适合研究时间序列聚类算法鲁棒性与诊断方法缺陷
时间序列聚类有望在医疗、金融、工业系统等领域揭示隐藏模式。然而,缺乏可靠的真值信息,使得研究者难以客观评估聚类质量,也无法判断不佳结果是源于数据本身无结构、算法局限还是评估方法不当,引发聚类是否“更像艺术而非科学”的质疑(Guyon et al., 2009)。为此,我们提出CSTS(Correlation Structures in Time Series),一个用于评估多变量时间序列中相关结构发现能力的合成基准。CSTS提供清晰的评估环境,可区分相关结构退化与聚类算法或验证方法的局限。贡献包括:(1) 具有明确相关结构、系统变化的数据条件、性能阈值和推荐评估协议的综合性基准;(2) 实证验证了降采样造成中等程度结构失真,而分布偏移和稀疏化影响极小;(3) 可扩展的数据生成框架,支持以结构为导向的聚类评估。案例研究揭示某算法对非正态分布的未记录敏感性,展示基准在精准诊断方法缺陷方面的实用性。CSTS推动了基于相关性的时序聚类的严格评估标准。
原文摘要 · Abstract (English)
Time series clustering promises to uncover hidden structural patterns in data with applications across healthcare, finance, industrial systems, and other critical domains. However, without validated ground truth information, researchers cannot objectively assess clustering quality or determine whether poor results stem from absent structures in the data, algorithmic limitations, or inappropriate validation methods, raising the question whether clustering is "more art than science" (Guyon et al., 2009). To address these challenges, we introduce CSTS (Correlation Structures in Time Series), a synthetic benchmark for evaluating the discovery of correlation structures in multivariate time series data. CSTS provides a clean benchmark that enables researchers to isolate and identify specific causes of clustering failures by differentiating between correlation structure deterioration and limitations of clustering algorithms and validation methods. Our contributions are: (1) a comprehensive benchmark for correlation structure discovery with distinct correlation structures, systematically varied data conditions, established performance thresholds, and recommended evaluation protocols; (2) empirical validation of correlation structure preservation showing moderate distortion from downsampling and minimal effects from distribution shifts and sparsification; and (3) an extensible data generation framework enabling structure-first clustering evaluation. A case study demonstrates CSTS's practical utility by identifying an algorithm's previously undocumented sensitivity to non-normal distributions, illustrating how the benchmark enables precise diagnosis of methodological limitations. CSTS advances rigorous evaluation standards for correlation-based time series clustering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。