用大模型分析时间序列数据,解决文本与数据间的跨模态鸿沟。
Towards Cross-Modality Modeling for Time Series Analytics: A Survey in the LLM Era
- 按文本类型分类现有方法,梳理跨模态融合路径。
- 实验证明特定文本+策略组合可显著提升时序分析效果。
- 适合关注大模型在时序数据应用的研究者与工程师。
边缘设备的普及催生了跨领域海量时间序列数据,推动了多种定制化分析方法的发展。近年来,大语言模型(LLMs)凭借其对序列数据的通用建模能力,成为时序分析的新范式。然而,由于LLMs在文本语料上预训练,缺乏对时间序列数据的原生优化,存在基础性的跨模态差距。近期多项工作致力于解决该问题。本文综述了基于大模型的时序数据分析跨模态建模最新进展。首先提出一个分类体系,将现有方法依据所用文本数据类型分为四类。接着总结关键跨模态策略,如对齐与融合,并讨论其在多类下游任务中的应用。此外,在多个应用领域的多模态数据集上进行实验,探究文本数据与跨模态策略的有效组合以增强时序分析性能。最后,指出若干未来研究方向。本综述面向对大模型驱动时序建模感兴趣的科研人员与实践者。
原文摘要 · Abstract (English)
The proliferation of edge devices has generated an unprecedented volume of time series data across different domains, motivating various well-customized methods. Recently, Large Language Models (LLMs) have emerged as a new paradigm for time series analytics by leveraging the shared sequential nature of textual data and time series. However, a fundamental cross-modality gap between time series and LLMs exists, as LLMs are pre-trained on textual corpora and are not inherently optimized for time series. Many recent proposals are designed to address this issue. In this survey, we provide an up-to-date overview of LLMs-based cross-modality modeling for time series analytics. We first introduce a taxonomy that classifies existing approaches into four groups based on the type of textual data employed for time series modeling. We then summarize key cross-modality strategies, e.g., alignment and fusion, and discuss their applications across a range of downstream tasks. Furthermore, we conduct experiments on multimodal datasets from different application domains to investigate effective combinations of textual data and cross-modality strategies for enhancing time series analytics. Finally, we suggest several promising directions for future research. This survey is designed for a range of professionals, researchers, and practitioners interested in LLM-based time series modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。