arXiv:2501.01832cs.CLcs.LG2025-01被引 13

用大模型生成时间序列描述,让数据会说话。

Time Series Language Model for Descriptive Caption Generation

  • 设计编码器-解码器结构,结合文本提示与时序数据生成描述
  • 通过上下文提示合成数据并去噪,解决标注数据少问题
  • 在多个数据集上超越现有方法,适合时序分析场景

自动生成可观测时间序列模式的自然语言描述,能提升可解释性、简化分析流程,并增强时序数据的跨领域应用价值。尽管预训练基础模型在自然语言处理(NLP)和计算机视觉(CV)中取得显著进展,其在时间序列分析中的应用仍受限于数据稀缺。虽然已有基于大语言模型(LLM)的时间序列预测方法,但时间序列字幕生成在LLM背景下仍研究不足。本文提出TSLM,一种专为时间序列字幕生成设计的新颖时序语言模型。TSLM采用编码器-解码器架构,利用文本提示和时序数据表示,捕捉多阶段细微时间模式,生成对输入时序数据的精确文字描述。针对时间序列字幕数据稀缺问题,TSLM首先通过上下文提示进行合成数据生成,再通过一种新颖的跨模态密集检索评分机制对生成数据进行去噪。在多个时间序列字幕生成数据集上的实验表明,TSLM在多个数据模态下显著优于现有最先进方法。

原文摘要 · Abstract (English)

The automatic generation of representative natural language descriptions for observable patterns in time series data enhances interpretability, simplifies analysis and increases cross-domain utility of temporal data. While pre-trained foundation models have made considerable progress in natural language processing (NLP) and computer vision (CV), their application to time series analysis has been hindered by data scarcity. Although several large language model (LLM)-based methods have been proposed for time series forecasting, time series captioning is under-explored in the context of LLMs. In this paper, we introduce TSLM, a novel time series language model designed specifically for time series captioning. TSLM operates as an encoder-decoder model, leveraging both text prompts and time series data representations to capture subtle temporal patterns across multiple phases and generate precise textual descriptions of time series inputs. TSLM addresses the data scarcity problem in time series captioning by first leveraging an in-context prompting synthetic data generation, and second denoising the generated data via a novel cross-modal dense retrieval scoring applied to time series-caption pairs. Experimental findings on various time series captioning datasets demonstrate that TSLM outperforms existing state-of-the-art approaches from multiple data modalities by a significant margin.

时序生成大模型字幕生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。