构建平衡数据集提升时间序列预测模型泛化能力
BLAST: Balanced Sampling Time Series Corpus for Universal Forecasting Models
- 基于网格采样与混洗技术,实现模式导向的均衡数据覆盖
- 包含3210亿条观测数据,显著提升模型训练效率与性能
- 适合需要高效通用预测模型的研究者与工业应用
通用时间序列预测模型的兴起推动了跨领域零样本预测的发展,但训练数据多样性这一关键因素仍缺乏深入研究。现有大规模时间序列数据集常存在固有偏差和分布不均问题,导致模型性能与泛化能力受限。为此,我们提出BLAST,一种通过均衡采样策略增强数据多样性的预训练语料库。BLAST整合了来自公开数据集的3210亿条观测,并利用全面的统计指标刻画时间序列模式;通过基于网格的分区方法对数据进行隐式聚类;结合网格采样与网格混洗技术,确保对多样化模式的均衡代表性覆盖。实验表明,基于BLAST预训练的模型在极低计算资源和训练词元消耗下达到顶尖性能。结果凸显数据多样性在提升通用预测任务训练效率与模型表现中的核心作用。
原文摘要 · Abstract (English)
The advent of universal time series forecasting models has revolutionized zero-shot forecasting across diverse domains, yet the critical role of data diversity in training these models remains underexplored. Existing large-scale time series datasets often suffer from inherent biases and imbalanced distributions, leading to suboptimal model performance and generalization. To address this gap, we introduce BLAST, a novel pre-training corpus designed to enhance data diversity through a balanced sampling strategy. First, BLAST incorporates 321 billion observations from publicly available datasets and employs a comprehensive suite of statistical metrics to characterize time series patterns. Then, to facilitate pattern-oriented sampling, the data is implicitly clustered using grid-based partitioning. Furthermore, by integrating grid sampling and grid mixup techniques, BLAST ensures a balanced and representative coverage of diverse patterns. Experimental results demonstrate that models pre-trained on BLAST achieve state-of-the-art performance with a fraction of the computational resources and training tokens required by existing methods. Our findings highlight the pivotal role of data diversity in improving both training efficiency and model performance for the universal forecasting task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。