arXiv:2605.31580cs.LG2026-05中稿 · ICML

用文本描述提升传感器时间序列的表示能力,让不同设备数据更易通用。

Giving Sensors a Voice: Multimodal JEPA for Semantic Time-Series Embeddings

论文配图:Giving Sensors a Voice: Multimodal JEPA for Semantic Time-Series Embeddings
图 1 · 摘自论文原文
  • 通过文本描述建模通道间关系,实现通道顺序不变的嵌入
  • 仅用线性探测器在异常检测、分类等任务上表现优异
  • 适合跨数据集、多传感器场景的通用时序建模

基于Transformer的架构推动了语言和视觉序列建模的发展,但异构多变量时间序列的通用表示学习仍不充分。我们提出CHRM(Channel-Aware Representation Model),将通道级文本描述融入对通道顺序不变的Transformer编码器中。模型采用联合嵌入预测架构(JEPA)和一种促进信息丰富、时间稳定的嵌入新损失函数;潜在空间预测增强对传感器噪声的鲁棒性,而描述感知门控则通过学习通道间关系提供可解释性。在异常检测、分类及短/长期预测任务中,所学嵌入仅用线性探测器即取得优异性能。表现主要由JEPA目标和条件架构驱动,文本描述作为通道标识符支持跨数据集泛化。

原文摘要 · Abstract (English)

Transformer-based architectures have advanced sequence modeling in language and vision, yet general-purpose representation learning for heterogeneous multivariate time series remains underexplored. We introduce CHARM (Channel-Aware Representation Model), which incorporates channel-level textual descriptions into a Transformer encoder equivariant to channel order. CHARM is trained with a Joint Embedding Predictive Architecture (JEPA) and a novel loss promoting informative, temporally stable embeddings; latent-space prediction encourages robustness to sensor noise while description-aware gating provides interpretability through learned inter-channel relationships. Across anomaly detection, classification, and short- and long-term forecasting, the learned embeddings achieve strong performance using only a linear probe. Performance is driven primarily by the JEPA objective and conditioning architecture, with text descriptions serving as channel identifiers for cross-dataset generalization.

时序建模多模态嵌入学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。