arXiv:2602.17149cs.LGcs.AI2026-02中稿 · ICML被引 2

首个统一时序理解与生成的视觉中心框架,提升精度与语义一致性。

TimeOmni-VL: Unified Models for Time Series Understanding and Generation

  • 通过双向图像-时序映射实现近无损转换
  • 理解任务引导生成,显著提升数值精度
  • 适合时序建模与多模态融合研究者

当前时序建模在数值生成与语义理解间存在明显分裂:生成模型常依赖表面模式匹配,而理解型模型难以输出高保真数值。尽管统一多模态模型(UMMs)已在视觉领域取得突破,但其在时序领域的潜力尚未被挖掘。本文提出TimeOmni-VL,首个以视觉为中心的统一框架,包含两项核心创新:(1) 时序与图像间的保真双向映射(Bi-TSI),显著提升时序转图像(TS2I)与图像转时序(I2TS)的转换质量;(2) 基于理解的生成机制。我们构建了包含六个理解任务与两个生成任务的全新数据集TSUMM-Suite,结合校准的思维链(Chain-of-Thought),首次将时序理解作为显式控制信号用于高保真生成。实验表明,该统一方法在语义理解与数值精度上均有显著提升,为多模态时序建模树立新标杆。

原文摘要 · Abstract (English)

Recent time series modeling faces a sharp divide between numerical generation and semantic understanding, with research showing that generation models often rely on superficial pattern matching, while understanding-oriented models struggle with high-fidelity numerical output. Although unified multimodal models (UMMs) have bridged this gap in vision, their potential for time series remains untapped. We propose TimeOmni-VL, the first vision-centric framework that unifies time series understanding and generation through two key innovations: (1) Fidelity-preserving bidirectional mapping between time series and images (Bi-TSI), which advances Time Series-to-Image (TS2I) and Image-to-Time Series (I2TS) conversions to ensure near-lossless transformations. (2) Understanding-guided generation. We introduce TSUMM-Suite, a novel dataset consisting of six understanding tasks rooted in time series analytics and coupled with two generation tasks. With a calibrated Chain-of-Thought, TimeOmni-VL is the first to leverage time series understanding as an explicit control signal for high-fidelity generation. Experiments confirm that this unified approach significantly improves semantic understanding and numerical precision, establishing a new frontier for multimodal time series modeling.

时序建模多模态生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。