提出ADAPT框架,实现多数据集混合训练的时序预训练。
ADAPTive Input Training for Many-to-One Pre-Training on Time-Series Classification

- 设计自适应输入机制,统一不同尺寸和通道数的时序数据。
- 在162个时序分类数据集上训练,达成新基准性能。
- 适合构建通用时序基础模型的研究者与工业应用开发者。
近期时序模型研究利用自监督学习来捕捉有意义的特征与模式,以提升下游任务表现并泛化到未见模态。尽管此类预训练方法在单源多目标场景中表现良好,但在预训练阶段加入更多数据集时,泛化能力显著下降。这成为构建时序领域基础模型的核心挑战。为此,本文提出一种名为ADAPT的新预训练范式,可高效对齐时序数据的物理属性,支持跨数据集混合批次训练,即使面对输入尺寸和通道维度差异极大的数据也有效。我们在162个时序分类数据集上进行训练,显著超越现有基准性能,成功实现多源数据同时预训练,为构建通用时序基础模型提供了关键支撑。
原文摘要 · Abstract (English)
Recent work on time-series models has leveraged self-supervised training to learn meaningful features and patterns in order to improve performance on downstream tasks and generalize to unseen modalities. While these pretraining methods have shown great promise in one-to-many scenarios, where a model is pre-trained on one dataset and fine-tuned on a downstream dataset, they have struggled to generalize to new datasets when more datasets are added during pre-training. This is a fundamental challenge in building foundation models for time-series data, as it limits the ability to develop models that can learn from a large variety of diverse datasets available. To address this challenge, we present a new pre-training paradigm for time-series data called ADAPT, which can efficiently align the physical properties of data in the time-series domain, enabling mixed-batch pre-training despite the extreme discrepancies in the input sizes and channel dimensions of pre-training data. We trained on 162 time-series classification datasets and set new state-of-the-art performance for classification benchmarks. We successfully train a model within the time-series domain on a wide range of datasets simultaneously, which is a major building block for building generalist foundation models in time-series domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。