用符号化数据生成解决时间序列模型训练数据少的问题
Synthetic Series-Symbol Data Generation for Time Series Foundation Models
- 基于动态系统理论设计符号化数据生成机制
- 新模型SymTime在5项任务上媲美真实数据预训练模型
- 适合需要小样本训练的时间序列研究者
时间序列分析的基础模型受到训练数据稀缺和不平衡的制约。受复杂动态系统理论启发,我们设计了一种序列-符号数据生成机制,可无限制生成高质量的时间序列数据及其对应符号表达式。为利用具有强相关性的序列-符号数据对,我们提出了SymTime,一个通过符号信息增强时间序列表征的预训练基础模型。在微调下游任务时,SymTime在五个主要时间序列分析任务中表现优异,性能媲美在真实世界数据上预训练的基础模型。该方法凸显了序列-符号数据生成与预训练机制在克服数据稀缺、提升任务性能方面的潜力。代码已开源:https://github.com/wwhenxuan/SymTime。
原文摘要 · Abstract (English)
Foundation models for time series analysis (TSA) have attracted significant attention. However, challenges such as training data scarcity and imbalance continue to hinder their development. Inspired by complex dynamic system theories, we design a series-symbol data generation mechanism, enabling the unrestricted creation of high-quality time series data paired with corresponding symbolic expressions. To leverage series-symbol data pairs with strong correlations, we develop SymTime, a pre-trained foundation model for enhancing time series representation using symbolic information. SymTime demonstrates competitive performance across five major TSA tasks when fine-tunes with downstream tasks, rivaling foundation models pre-trained on real-world datasets. This approach underscores the potential of series-symbol data generation and pretraining mechanisms in overcoming data scarcity and enhancing task performance. The code is available at https://github.com/wwhenxuan/SymTime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。