arXiv:2502.15122cs.LG2025-02被引 6

构建大规模时间序列分类数据集,推动模型可扩展性研究

MONSTER: Monash Scalable Time Series Evaluation Repository

  • 构建大型时间序列分类数据集集合,突破传统小规模基准限制
  • 数据集规模远超UCR/UEA(中位数达数百至数千样本)
  • 适合关注模型效率与大数据学习的算法研究者

我们提出MONSTER——莫纳什大学可扩展时间序列评估库,一个用于时间序列分类的大规模数据集集合。当前时间序列分类领域依赖的UCR和UEA基准数据集规模较小,中位数分别为217和255个样本。这导致研究过度聚焦于低方差、小规模优化模型,忽视了计算可扩展性等实际挑战。MONSTER通过引入更大规模的数据集,旨在拓展研究边界,激发在大规模数据上高效学习的理论与实践进展。

原文摘要 · Abstract (English)

We introduce MONSTER-the MONash Scalable Time Series Evaluation Repository-a collection of large datasets for time series classification. The field of time series classification has benefitted from common benchmarks set by the UCR and UEA time series classification repositories. However, the datasets in these benchmarks are small, with median sizes of 217 and 255 examples, respectively. In consequence they favour a narrow subspace of models that are optimised to achieve low classification error on a wide variety of smaller datasets, that is, models that minimise variance, and give little weight to computational issues such as scalability. Our hope is to diversify the field by introducing benchmarks using larger datasets. We believe that there is enormous potential for new progress in the field by engaging with the theoretical and practical challenges of learning effectively from larger quantities of data.

时间序列数据集可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。