arXiv:2511.19856cs.CV2025-11

让时间序列直接生成图像,实现时序与视觉的语义对齐。

Temporal-Visual Semantic Alignment: A Unified Architecture for Transferring Spatial Priors from Vision Models to Zero-Shot Temporal Tasks

  • 通过双自编码器和共享量化器学习跨模态表示,再用投影对齐时序与视觉特征。
  • 在零样本时序任务上表现优异,生成图像质量高且能体现时间波动模式。
  • 适合需要从时间数据生成视觉内容的研究者,如金融、医疗可视化场景。

大型多模态模型在文本与图像对齐与生成方面取得显著进展,但将非视觉的连续序列作为高质量图像生成条件的潜力尚未被充分探索。现有将序列转为“伪图像”进行时序预测的方法,未能建立语义层面的对齐。本文提出 TimeArtist,一个时序-视觉转换框架,首次实现时间序列波动与视觉概念之间的语义级对齐。其开创性地采用“预热对齐”范式:首先在大规模数据集上通过双自编码器和共享量化器进行自监督训练,学习模态共享表示;随后冻结编码器与量化器,引入投影层在表征层面对齐时序与视觉样本。TimeArtist 构建了一个通用的跨模态框架,可直接从时间序列生成高质量、多样化的图像,同时捕捉时间波动模式以实现风格迁移。大量实验表明,该方法在图像生成指标上表现良好,并在零样本时序任务中取得更优结果。本工作建立了跨模态生成的新范式,弥合了时序动态与视觉语义之间的鸿沟。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) have achieved remarkable progress in aligning and generating content across text and image modalities. However, the potential of using non-visual, continuous sequential, as a conditioning signal for high-fidelity image generation remains largely unexplored. Furthermore, existing methods that convert series into "pseudo-images" for temporal forecasting fail to establish semantic-level alignment. In this paper, we propose TimeArtist, a temporal-visual conversion framework that pioneers semantic-level alignment between time series fluctuations and visual concepts. It pioneers a "warmup-align" paradigm: first, a dual-autoencoder and shared quantizer are self-supervised trained on large-scale datasets to learn modality-shared representations. Then, the encoders and quantizer are frozen, and a projection is introduced to align temporal and visual samples at the representation level. TimeArtist establishes a versatile cross-modal framework, enabling high-quality, diverse image generation directly from time series, while capturing temporal fluctuation patterns to render images as styles transfer. Extensive experiments show that TimeArtist achieves satisfactory performance in image generation metrics, while also attaining superior results in zero-shot temporal tasks. Our work establishes a new paradigm for cross-modal generation, bridging the gap between temporal dynamics and visual semantics.

跨模态生成时间序列图像生成语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。