用大模型理解生成科学时间序列,突破传统处理方式局限。
SciTS: Scientific Time Series Understanding and Generation with LLMs
- 将时间序列直接输入大模型,避免转文本或图像带来的信息损失。
- 在43个任务上测试17种模型,发现通用大模型泛化能力更强。
- 提出TimeOmni框架,支持复杂科学时序数据的理解与生成。
大语言模型(LLMs)在科学推理中的潜力日益受到关注。时间序列作为科学数据的基本模态,当前多模态大模型常将其编码为文本或转换为图像,但此类方法在处理长序列或高精度数值时存在不足。现有统一的时间序列模型通常专注于预测或分析,对非周期性、异构科学信号的效能尚不明确。为此,我们构建了涵盖12个科学领域、43项任务的SciTS基准,包含超过5万条样本,涵盖从10⁰到10⁷长度、最高达10 MHz频率的单变量与多变量信号。我们评估了17种模型,包括纯文本大模型、多模态大模型和专用时间序列模型,发现通用大模型在泛化能力上优于专用模型,而文本或图像表示会因序列过长或精度丢失导致性能下降。为此,我们提出TimeOmni框架,使大模型具备理解与生成时间序列的能力,且兼容通用训练流程。该工作填补了科学时间序列专用基准与建模范式的空白,推动大模型在复杂时序科学数据上的应用。
原文摘要 · Abstract (English)
The scientific reasoning ability of large language models (LLMs) has recently attracted significant attention. Time series, as a fundamental modality in scientific data, presents unique challenges that are often overlooked in current multimodal LLMs, which either encode numerical sequences as text or convert them into images. Such approaches may be insufficient for comprehensive scientific time series understanding and generation. Existing unified time series models typically specialise in either forecasting or analysis, and their effectiveness on non-periodic, heterogeneous scientific signals remains unclear. To address these gaps, we introduce SciTS, a benchmark spanning 12 scientific domains and 43 tasks, with over 50k+ instances, both univariate and multivariate signals ranging from $10^0$ to $10^7$ in length and up to 10~MHz in frequency. We benchmark 17 models, including text-only LLMs, multimodal LLMs, and unified time series models, and find that general-purpose LLMs exhibit stronger generalisability than specialised time series models, while representing time series as text or images limits their performance due to excessively long sequences and loss of numerical precision, respectively. We then introduce TimeOmni, a framework that equips LLMs with the ability to understand and generate time series while remaining compatible with general-purpose LLM training. This work fills a gap in both dedicated benchmarks and modelling frameworks for scientific time series, paving the way for LLMs to understand and generate complex temporal scientific data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。