通过频谱量化实现时间序列跨域迁移,无需标注数据也能高效适配新任务。
A Wave is Worth 100 Words: Investigating Cross-Domain Transferability in Time Series
- 将多源时间序列映射到统一频谱隐空间,实现跨域知识共享。
- 在87.5%的任务上表现最优,平均性能提升34.7%。
- 适用于零样本、少样本场景,兼容主流模型架构。
时间序列分析是基础的数据挖掘任务,基于经验风险最小化的监督学习方法在特定任务和数据集上已证明有效。然而,高质量标注数据获取成本高,大量未标注序列数据被闲置。由于不同领域间分布差异及任务模式多样,时间序列的跨域多任务迁移仍面临重大挑战。为此,本文提出一种基于频谱量化(WQ4TS)的新颖跨域预训练方法,可与任意先进时间序列模型结合,应用于多个下游任务。具体而言,将不同领域的时序数据转换至统一的频谱隐空间,使模型能直接从该公共空间中学习各领域的时间模式知识,并用于下游任务推理,从而缓解异构跨域迁移难题。频谱隐空间的建立带来三大优势:具备跨域迁移能力,可在无先验知识条件下适应零样本与少样本场景;构建通用兼容的跨域迁移框架,无需修改现有模型结构;具备强健建模能力,在多个下游任务中达到最先进水平。为验证所提方法的有效性,我们在三个重要任务(预测、缺失值填补、分类)上开展广泛实验,并模拟三种真实场景:全数据、少样本、零样本。结果表明,WQ4TS在所有任务中87.5%的表现最优,整体指标平均提升达34.7%。
原文摘要 · Abstract (English)
Time series analysis is a fundamental data mining task that supervised training methods based on empirical risk minimization have proven their effectiveness on specific tasks and datasets. However, the acquisition of well-annotated data is costly and a large amount of unlabeled series data is under-utilized. Due to distributional shifts across various domains and different patterns of interest across multiple tasks. The problem of cross-domain multi-task migration of time series remains a significant challenge. To address these problems, this paper proposes a novel cross-domain pretraining method based on Wave Quantization (termed as WQ4TS), which can be combined with any advanced time series model and applied to multiple downstream tasks. Specifically, we transfer the time series data from different domains into a common spectral latent space, and enable the model to learn the temporal pattern knowledge of different domains directly from the common space and utilize it for the inference of downstream tasks, thereby mitigating the challenge of heterogeneous cross-domains migration. The establishment of spectral latent space brings at least three benefits, cross-domain migration capability thus adapting to zero- and few-shot scenarios without relying on priori knowledge of the dataset, general compatible cross-domain migration framework without changing the existing model structure, and robust modeling capability thus achieving SOTA results in multiple downstream tasks. To demonstrate the effectiveness of the proposed approach, we conduct extensive experiments including three important tasks: forecasting, imputation, and classification. And three common real-world data scenarios are simulated: full-data, few-shot, and zero-shot. The proposed WQ4TS achieves the best performance on 87.5% of all tasks, and the average improvement of the metrics on all the tasks is up to 34.7%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。