构建大规模真实世界多变量时间序列数据集,验证真实数据对模型泛化能力的提升。
RMISC: A Large-scale Real-world Multivariate Corpus for Time Series Foundation Models

- 建立包含1420亿时间点的大型真实世界多变量时间序列数据集RMISC。
- 实测表明,用真实数据训练的模型在零样本任务中表现更优。
- 适合时间序列建模、金融/工业预测等领域研究者参考。
近年来,多变量时间序列基础模型(TSFMs)兴起,展现出优异的零样本泛化能力。当前主流模型多基于易于扩展的合成多变量数据预训练,但这类数据可能无法捕捉真实时间序列中的复杂时序动态与变量间关系。这引出一个关键问题:使用真实世界语料训练的先进TSFMs是否优于合成数据训练的模型?为回答此问题,我们构建了RMISC语料库——一个大规模、高质量、公开可访问的真实世界多变量时间序列档案,涵盖约200个数据集和1420亿个时间点,覆盖多种领域。我们进一步在单变量、合成多变量和真实多变量数据上预训练四种先进TSFMs,并在标准的分布内与分布外基准上评估其零样本泛化能力。实验结果表明,引入真实世界多变量数据显著提升了单变量与多变量TSFMs的泛化性能。这些结果深化了对真实世界多变量数据如何促进更强TSFMs发展的理解。
原文摘要 · Abstract (English)
Recent years have witnessed the emergence of multivariate modeling using time series foundation models (TSFMs), which achieve advanced zero-shot generalization. Modern multivariate TSFMs are predominantly pretrained on multivariate synthetic data, which is easier to scale but may fail to capture the complex temporal dynamics and cross-variable relationships present in real-world time series. This raises a key question: Whether and to what extent the leading TSFMs trained with the real-world corpus perform better than those trained with synthetic data? To answer this, we establish the RMISC corpus, a considerably large-scale, high-quality, openly accessible, real-world, and multivariate time series archive that contains around 200 datasets and 142 billion time points across diverse domains. Furthermore, we pretrain four advanced TSFMs on univariate, synthetic multivariate, and real-world multivariate data and evaluate their zero-shot generalization capabilities on standard in-distribution and out-of-distribution benchmarks. Experimental results show that incorporating real-world multivariate data predominantly improves the generalization performance for both univariate and multivariate TSFMs. These results provide a deeper understanding of how real-world multivariate data contributes to the development of stronger TSFMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。