arXiv:2603.20352cs.LG2026-03被引 2

扩充了多变量时间序列分类数据集,总量超140个,支持更广泛研究。

The Multiverse of Time Series Machine Learning: an Archive for Multivariate Time Series Classification

  • 将原有30个数据集扩展至147个,含缺失值与不等长序列预处理版本。
  • 新增117个分类任务,覆盖多领域,建立统一开源资源库。
  • 提供可复现的框架和基准测试,适合新手快速上手与对比实验。

时间序列机器学习(TSML)研究不断增长,其发展很大程度依赖于基准数据集。2018年发布的UEA多变量时间序列分类数据集档案,已成为数百篇论文引用的核心资源。本文大幅扩展该档案,将数据集数量从30个增至133个,新增预处理版本以涵盖存在缺失值或不等长序列的数据,总规模达147个。为体现其多样性,新档案更名为「Multiverse archive」,整合多个来源的独立数据集与集合,形成统一仓库。鉴于全量实验计算成本高,我们推荐使用子集「MV-core」进行初步探索。同时提供详细使用指南与主流算法基准评估,确立未来研究性能参照标准。配套建立专用仓库,支持可复现性,兼容scikit-learn,并提供交互式界面查看已发表结果。

原文摘要 · Abstract (English)

Time series machine learning (TSML) is a growing research field that spans a wide range of tasks. The popularity of established tasks such as classification, clustering, and extrinsic regression has, in part, been driven by the availability of benchmark datasets. An archive of 30 multivariate time series classification datasets, introduced in 2018 and commonly known as the UEA archive, has since become an essential resource cited in hundreds of publications. We present a substantial expansion of this archive that more than quadruples its size, from 30 to 133 classification problems. We also release preprocessed versions of datasets containing missing values or unequal length series, bringing the total number of datasets to 147. Reflecting the growth of the archive and the broader community, we rebrand it as the Multiverse archive to capture its diversity of domains. The Multiverse archive includes datasets from multiple sources, consolidating other collections and standalone datasets into a single, unified repository. Recognising that running experiments across the full archive is computationally demanding, we recommend a subset of the full archive called Multiverse-core (MV-core) for initial exploration. To support researchers in using the new archive, we provide detailed guidance and a baseline evaluation of established and recent classification algorithms, establishing performance benchmarks for future research. We have created a dedicated repository for the Multiverse archive that provides a common aeon and scikit-learn compatible framework for reproducibility, an extensive record of published results, and an interactive interface to explore the results.

时间序列数据集分类开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。