arXiv:2502.15466cs.LGcs.AI2025-02被引 3

用符号化数据生成解决时间序列数据稀缺问题,提升模型性能。

Mitigating Data Scarcity in Time Series Analysis: A Foundation Model with Series-Symbol Data Generation

  • 通过符号表达生成高质量时间序列与对应符号表示
  • 在5个主流任务上表现媲美真实数据预训练模型
  • 适合需要小样本学习的时间序列研究者

时间序列分析(TSA)的基础模型受到广泛关注,但数据稀缺和数据不平衡仍是主要挑战。为此,本文提出以符号表达建模复杂系统,作为时间序列的语义描述。基于此思想,构建了系列-符号(S2)双模态数据生成机制,可无限制生成高质量时间序列及其对应符号表示。利用S2数据集,我们开发了SymTime基础模型。在下游任务微调后,SymTime在五个主要TSA任务中表现优异,性能媲美在真实世界数据上预训练的模型。该方法展示了双模态数据生成与预训练机制在克服数据稀缺、提升任务表现方面的潜力。

原文摘要 · Abstract (English)

Foundation models for time series analysis (TSA) have attracted significant attention. However, challenges such as data scarcity and data imbalance continue to hinder their development. To address this, we consider modeling complex systems through symbolic expressions that serve as semantic descriptors of time series. Building on this concept, we introduce a series-symbol (S2) dual-modulity data generation mechanism, enabling the unrestricted creation of high-quality time series data paired with corresponding symbolic representations. Leveraging the S2 dataset, we develop SymTime, a pre-trained foundation model for TSA. SymTime demonstrates competitive performance across five major TSA tasks when fine-tuned with downstream task, rivaling foundation models pre-trained on real-world datasets. This approach underscores the potential of dual-modality data generation and pretraining mechanisms in overcoming data scarcity and enhancing task performance.

时间序列基础模型数据生成符号学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。