用频谱当桥梁,让时间序列、文本和图像一起学工业设备健康预测。
VLT: A Vision-Language-Time Series Multimodal Foundation Model for Industrial Intelligence

- 用频谱图做视觉桥梁,连通时间信号与文本语义。
- 在少样本、噪声、缺模态下仍表现更好,准确率超现有方法。
- 适合工业智能、设备故障预测场景的工程师和研究者。
工业时间序列是保障航空发动机等设备可靠性和安全性的核心数据,支撑着故障预测与健康管理(PHM)。然而,现有方法多为单模态建模,难以适应复杂场景下的泛化需求。尽管大语言模型(LLMs)为多模态学习带来新机遇,如何融合连续时间信号与离散文本语义仍是未解难题。为此,我们提出VLT,一种联合建模时间序列、频谱视觉表征与文本知识的多模态基础模型。关键思想是将频谱作为连接连续时序信号与离散语义的视觉桥梁。具体地,设计了时序感知的专家混合(Time-MoE)以捕捉异构时序动态;引入频谱-文本增强学习器,在共享表示空间中联合建模谱特征与语义特征。此外,提出以时间为中心的梯度对齐机制,通过梯度归一化与可靠性感知的动态重加权缓解跨模态优化冲突。在多个工业数据集上的实验表明,VLT显著优于当前最优方法,在少样本、含噪、不完整模态等挑战性设置下展现出更强鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Industrial time series serve as the foundation for Prognostics and Health Management (PHM) to ensure the reliability and safety of industrial equipment such as aero-engines. However, existing approaches are typically limited to single-modality modeling, which restricts their generalization in complex scenarios. Although recent advances in large language models (LLMs) provide new opportunities for multimodal learning, bridging continuous time-series signals and discrete textual semantics remains an open challenge. To this end, we propose VLT, a multimodal foundation model that jointly models time-series, frequency-spectrum visual representations, and textual knowledge. A key insight is to utilize the frequency spectrum as a visual bridge to connect continuous temporal signals with discrete semantics. Specifically, a Time-aware Mixture-of-Experts (Time-MoE) is designed to capture heterogeneous temporal dynamics, while a Frequency-Text Augmented Learner enables joint modeling of spectral and semantic features within a shared representation space. Furthermore, a time-centric gradient alignment mechanism is introduced to mitigate cross-modal optimization conflicts via gradient normalization and reliability-aware dynamic reweighting. Extensive experiments on multiple industrial datasets demonstrate that VLT outperforms state-of-the-art methods, achieving superior robustness and generalization under few-shot, noisy, and incomplete-modality settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。