通过联邦学习训练跨领域时间序列基础模型,解决数据异构难题
Federated Foundation Models on Heterogeneous Time Series
- 采用联邦框架,各机构本地训练专属模型保留数据特性
- 引入双端正则化机制,对齐不同领域间的共享知识
- 在预测、填补和异常检测任务中表现优于现有方法
从头训练具备强泛化能力的通用时间序列基础模型仍是开放挑战。现有方法多依赖跨领域时间序列数据融合以提取共享子序列作为Transformer模型的输入,但由于领域间显著的统计异质性,该方法在时间序列上效果不如文本和图像。为此,本文提出一种新型联邦学习方法FFTS:将每个数据持有机构视为独立客户端,在联邦设置下分别训练本地模型以保留数据集特有特征;同时引入客户端与服务器端的双重正则化机制,对齐来自不同领域的共享知识。在基准数据集上的大量实验表明,所学得的时间序列基础模型在跨领域分析任务(包括预测、填补和异常检测)中展现出优越的泛化能力。
原文摘要 · Abstract (English)
Training a general-purpose time series foundation models with robust generalization capabilities across diverse applications from scratch is still an open challenge. Efforts are primarily focused on fusing cross-domain time series datasets to extract shared subsequences as tokens for training models on Transformer architecture. However, due to significant statistical heterogeneity across domains, this cross-domain fusing approach doesn't work effectively as the same as fusing texts and images. To tackle this challenge, this paper proposes a novel federated learning approach to address the heterogeneity in time series foundation models training, namely FFTS. Specifically, each data-holding organization is treated as an independent client in a collaborative learning framework with federated settings, and then many client-specific local models will be trained to preserve the unique characteristics per dataset. Moreover, a new regularization mechanism will be applied to both client-side and server-side, thus to align the shared knowledge across heterogeneous datasets from different domains. Extensive experiments on benchmark datasets demonstrate the effectiveness of the proposed federated learning approach. The newly learned time series foundation models achieve superior generalization capabilities on cross-domain time series analysis tasks, including forecasting, imputation, and anomaly detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。