对比124个时间序列分类任务,发现特征集表现差异不大,但成分设计影响关键性能。
Statistical comparisons of time-series feature sets on classification tasks

- 采用归一化基准法,系统比较6个开源特征集与3个基础特征集
- 85.3%的对比结果为平局,tsfresh在29.03%对比中胜出
- 特定问题上特征组成或简单基线(傅里叶系数+分位数)更具优势
近年来,众多开源软件库被开发用于计算单变量时间序列的特征集。这些特征集在类型和数量上差异显著,其构建基于不同的学科视角来量化时间序列结构。然而,这些特征集在时间序列分类任务中的相对优劣仍缺乏深入研究。本文通过一种基于归一化的任务级基准方法,评估了六个开源特征集与三个基于分布和/或基本频谱结构的基准特征集在124个单变量时间序列分类问题上的表现,该方法相比以往基于排名的方法更准确地衡量算法间的相对优劣。尽管特征集在规模、组成和计算时间上差异巨大,整体表现却相对接近(85.3%的成对比较结果为平局),其中规模最大的特征集tsfresh表现出最强的整体性能(在所有成对比较中取得29.03%的胜利)。我们还识别出某些特定问题中特征集的特定构成会带来显著优势或劣势,并发现一些问题仅靠傅里叶系数和分位数等简单基线即可达到良好性能。结果表明,评估时间序列特征集时需考虑任务级表现,且特征构成是决定分类性能的关键因素。
原文摘要 · Abstract (English)
In recent years, numerous open-source software libraries have been developed for computing sets of features from univariate time series. The type and number of features vary across these feature sets, which have been constructed with varying disciplinary perspectives on quantifying structure in time-series data. To date, the relative strengths and weaknesses of these feature sets on time-series classification problems remains largely unexplored. Here we aimed to understand the relative performance of six open-source feature sets and three baseline feature sets (based on distributional and/or basic spectral structure) across 124 univariate time-series classification problems using a normalization-based approach to problem-level benchmarking that better indexes the relative strengths and weaknesses of different algorithms compared to prior rank-based approaches. Despite their dramatic differences in size, composition, and computation time, we found that feature sets performed relatively similarly overall (85.3% of pairwise comparisons resulted in ties), with the largest feature set, tsfresh, exhibiting the strongest overall performance (29.03% wins across all pairwise comparisons against other feature sets). We also highlighted specific problems on which the specific composition of a given feature set gave it a substantial performance advantage or disadvantage, and problems where simple baselines comprised of Fourier coefficients and quantiles were sufficient to achieve strong performance. Our results demonstrate the need to consider problem-level performance when benchmarking time-series feature sets, and highlight the importance of feature make-up in driving relative classification performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。