用合成数据解决时间序列模型训练数据难题
Empowering Time Series Analysis with Synthetic Data: A Survey and Outlook in the Era of Foundation Models
- 通过合成数据生成技术补充真实数据不足
- 支持基础模型预训练、微调与评估全流程
- 适合关注数据瓶颈的时间序列研究者
时间序列分析对理解复杂系统动态至关重要。近年来,任务无关的时间序列基础模型(TSFMs)和基于大语言模型的时间序列模型(TSLLMs)兴起,实现了泛化学习并融合上下文信息。然而,其成功依赖于大规模、多样化且高质量的数据集,而真实数据受法规、多样性、质量和数量限制难以构建。合成数据成为可行解决方案,可提供可扩展、无偏见、高质量的数据替代方案。本文全面综述了面向TSFMs和TSLLMs的合成数据技术,分析数据生成策略及其在模型预训练、微调和评估中的作用,并指出未来研究方向。
原文摘要 · Abstract (English)
Time series analysis is crucial for understanding dynamics of complex systems. Recent advances in foundation models have led to task-agnostic Time Series Foundation Models (TSFMs) and Large Language Model-based Time Series Models (TSLLMs), enabling generalized learning and integrating contextual information. However, their success depends on large, diverse, and high-quality datasets, which are challenging to build due to regulatory, diversity, quality, and quantity constraints. Synthetic data emerge as a viable solution, addressing these challenges by offering scalable, unbiased, and high-quality alternatives. This survey provides a comprehensive review of synthetic data for TSFMs and TSLLMs, analyzing data generation strategies, their role in model pretraining, fine-tuning, and evaluation, and identifying future research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。