arXiv:2506.22927cs.LG2025-06

用自然语言生成时间序列,让文本直接变数据。

Towards Time Series Generation Conditioned on Unstructured Natural Language

  • 用扩散模型与语言模型结合,从文本生成时间序列。
  • 在6.3万对数据上验证,能准确响应自然语言描述。
  • 适合需要自定义数据生成的研究者和开发者。

生成式人工智能已能生成图像、文本等多种数据,但时间序列生成仍较滞后,尽管其在金融、气候等领域的应用至关重要。本文提出一种新方法,通过结合扩散模型与语言模型,实现基于非结构化自然语言描述的时间序列生成。实验表明该方法可行,可支持定制化预测、数据增强、时序操控及迁移学习等应用。此外,我们构建并发布了一个新的公开数据集,包含63,010个时间序列-描述配对,推动该领域发展。

原文摘要 · Abstract (English)

Generative Artificial Intelligence (AI) has rapidly become a powerful tool, capable of generating various types of data, such as images and text. However, despite the significant advancement of generative AI, time series generative AI remains underdeveloped, even though the application of time series is essential in finance, climate, and numerous fields. In this research, we propose a novel method of generating time series conditioned on unstructured natural language descriptions. We use a diffusion model combined with a language model to generate time series from the text. Through the proposed method, we demonstrate that time series generation based on natural language is possible. The proposed method can provide various applications such as custom forecasting, time series manipulation, data augmentation, and transfer learning. Furthermore, we construct and propose a new public dataset for time series generation, consisting of 63,010 time series-description pairs.

时序生成语言控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。