让时间序列会读会写,用文本+数据联合预测未来
ChatTime: A Unified Multimodal Time Series Foundation Model Bridging Numerical and Textual Data

- 把时间序列当外语处理,统一建模数值与文本数据
- 零样本预测能力,跨场景适配且支持图文双向输出
- 适合需要融合专家经验与数据的金融、医疗等场景
人类专家通常结合数值与文本多模态信息分析时间序列。然而,大多数传统深度学习预测模型仅依赖单模态数值数据,使用固定长度窗口在单一数据集上训练和预测,无法适应不同场景。预训练大语言模型为时间序列分析带来新机遇。但现有方法或训练效率低,或无法处理文本信息,或缺乏零样本预测能力。本文创新性地将时间序列视为外语,构建了ChatTime——一个统一的时间序列与文本处理框架。作为开箱即用的多模态时间序列基础模型,ChatTime具备零样本预测能力,支持时间序列与文本的双向输入输出。我们设计了一系列实验验证其在多个任务和场景下的优越性能,并构建了四个多模态数据集以填补数据空白。实验结果证明了ChatTime的潜力与实用性。
原文摘要 · Abstract (English)
Human experts typically integrate numerical and textual multimodal information to analyze time series. However, most traditional deep learning predictors rely solely on unimodal numerical data, using a fixed-length window for training and prediction on a single dataset, and cannot adapt to different scenarios. The powered pre-trained large language model has introduced new opportunities for time series analysis. Yet, existing methods are either inefficient in training, incapable of handling textual information, or lack zero-shot forecasting capability. In this paper, we innovatively model time series as a foreign language and construct ChatTime, a unified framework for time series and text processing. As an out-of-the-box multimodal time series foundation model, ChatTime provides zero-shot forecasting capability and supports bimodal input/output for both time series and text. We design a series of experiments to verify the superior performance of ChatTime across multiple tasks and scenarios, and create four multimodal datasets to address data gaps. The experimental results demonstrate the potential and utility of ChatTime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。