arXiv:2505.02417cs.LGcs.AI2025-05IJCAI被引 22

用扩散模型实现任意长度高分辨率时序数据的文本生成。

T2S: High-resolution Time Series Generation with Text-to-Series Diffusion Models

  • 按点、片段、实例三层次构建时序描述,提升泛化能力。
  • 在13个跨领域数据集上达当前最优,支持任意长度生成。
  • 适合需要高质量时序数据合成的研究者与工程师。

文本到时序数据生成在应对数据稀疏、不平衡及多模态时序数据集稀缺等问题上具有重要潜力。尽管扩散模型在文本到图像、音频等任务中表现卓越,但其在时序数据生成领域仍处于初期阶段。现有方法存在两大瓶颈:(1)缺乏对通用时序描述的系统探索,现有描述常具领域特异性且泛化能力差;(2)无法生成任意长度的时序序列,限制实际应用。本文首次将时序描述分为点级、片段级和实例级三类,并构建了一个包含超过60万条高分辨率时序-文本对的新片段级数据集。提出Text-to-Series(T2S)框架,基于扩散模型实现跨领域无偏时序生成。T2S采用可变长度的变分自编码器,将不同长度的时序数据统一编码为一致的潜在表示;通过流匹配对齐文本与潜空间表示,并使用扩散Transformer作为去噪器。在多长度交替训练策略下,T2S可生成任意指定长度的序列。大量实验表明,T2S在13个涵盖12个领域的数据集上均达到领先性能。

原文摘要 · Abstract (English)

Text-to-Time Series generation holds significant potential to address challenges such as data sparsity, imbalance, and limited availability of multimodal time series datasets across domains. While diffusion models have achieved remarkable success in Text-to-X (e.g., vision and audio data) generation, their use in time series generation remains in its nascent stages. Existing approaches face two critical limitations: (1) the lack of systematic exploration of general-proposed time series captions, which are often domain-specific and struggle with generalization; and (2) the inability to generate time series of arbitrary lengths, limiting their applicability to real-world scenarios. In this work, we first categorize time series captions into three levels: point-level, fragment-level, and instance-level. Additionally, we introduce a new fragment-level dataset containing over 600,000 high-resolution time series-text pairs. Second, we propose Text-to-Series (T2S), a diffusion-based framework that bridges the gap between natural language and time series in a domain-agnostic manner. T2S employs a length-adaptive variational autoencoder to encode time series of varying lengths into consistent latent embeddings. On top of that, T2S effectively aligns textual representations with latent embeddings by utilizing Flow Matching and employing Diffusion Transformer as the denoiser. We train T2S in an interleaved paradigm across multiple lengths, allowing it to generate sequences of any desired length. Extensive evaluations demonstrate that T2S achieves state-of-the-art performance across 13 datasets spanning 12 domains.

时序生成扩散模型文本生成数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。