为时间序列设计可学习的伪词嵌入,提升大模型预测性能
VITRO: Vocabulary Inversion for Time-series Representation Optimization
- 用文本反演思路构建专用于时间序列的可学习伪词
- 在多数数据集上实现长期预测的最先进效果
- 适合需要高精度时序建模的研究者和工程师
尽管大语言模型在文本处理与生成方面表现出色,但其预训练词汇难以捕捉时间序列中细微的动态与模式。自然语言标记的离散符号特性与时间序列的连续数值特性不匹配。为此,我们提出VITRO,借鉴视觉-语言领域的文本反演优化方法,学习针对特定数据集的时间序列专用伪词嵌入,弥合自然语言的离散语义与时间序列的连续数值之间的鸿沟。实验表明,可学习的时间序列特有伪词嵌入比通用语言模型词汇更能有效表示时间序列数据,在多数数据集上,VITRO增强的方法实现了长期预测的最先进性能。
原文摘要 · Abstract (English)
Although LLMs have demonstrated remarkable capabilities in processing and generating textual data, their pre-trained vocabularies are ill-suited for capturing the nuanced temporal dynamics and patterns inherent in time series. The discrete, symbolic nature of natural language tokens, which these vocabularies are designed to represent, does not align well with the continuous, numerical nature of time series data. To address this fundamental limitation, we propose VITRO. Our method adapts textual inversion optimization from the vision-language domain in order to learn a new time series per-dataset vocabulary that bridges the gap between the discrete, semantic nature of natural language and the continuous, numerical nature of time series data. We show that learnable time series-specific pseudo-word embeddings represent time series data better than existing general language model vocabularies, with VITRO-enhanced methods achieving state-of-the-art performance in long-term forecasting across most datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。