arXiv:2603.07572cs.LG2026-03被引 1

用多模态大模型统一分析工业时序数据,提升预测准确率。

TS-MLLM: A Multi-Modal Large Language Model-based Framework for Industrial Time-Series Big Data Analysis

  • 融合时序、频域图像和文本知识的多模态建模方法
  • 在少样本和复杂场景下性能显著优于现有方法
  • 适合工业设备健康监测与故障预测场景

工业时序大数据的精准分析对设备健康管理至关重要。尽管大语言模型在时序分析中展现出潜力,但现有方法多局限于单模态适配,未能充分利用时序信号、频域视觉表示与文本知识之间的互补性。本文提出TS-MLLM,一种统一的多模态大模型框架,联合建模时序信号、频域图像与文本领域知识。首先设计工业时序块建模分支以捕捉长程时序动态;为引入跨模态先验,提出谱感知视觉-语言模型适配(SVLMA)机制,使模型内化频域模式与语义上下文;进一步设计以时序为中心的多模态注意力融合(TMAF)机制,利用时序特征作为查询主动检索相关视觉与文本线索,实现深度跨模态对齐。在多个工业基准上的实验表明,TS-MLLM显著优于现有最优方法,尤其在少样本和复杂场景下表现突出,验证了其在工业时序预测中的卓越鲁棒性、效率与泛化能力。

原文摘要 · Abstract (English)

Accurate analysis of industrial time-series big data is critical for the Prognostics and Health Management (PHM) of industrial equipment. While recent advancements in Large Language Models (LLMs) have shown promise in time-series analysis, existing methods typically focus on single-modality adaptations, failing to exploit the complementary nature of temporal signals, frequency-domain visual representations, and textual knowledge information. In this paper, we propose TS-MLLM, a unified multi-modal large language model framework designed to jointly model temporal signals, frequency-domain images, and textual domain knowledge. Specifically, we first develop an Industrial time-series Patch Modeling branch to capture long-range temporal dynamics. To integrate cross-modal priors, we introduce a Spectrum-aware Vision-Language Model Adaptation (SVLMA) mechanism that enables the model to internalize frequency-domain patterns and semantic context. Furthermore, a Temporal-centric Multi-modal Attention Fusion (TMAF) mechanism is designed to actively retrieve relevant visual and textual cues using temporal features as queries, ensuring deep cross-modal alignment. Extensive experiments on multiple industrial benchmarks demonstrate that TS-MLLM significantly outperforms state-of-the-art methods, particularly in few-shot and complex scenarios. The results validate our framework's superior robustness, efficiency, and generalization capabilities for industrial time-series prediction.

时序分析多模态大模型工业智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。