arXiv:2602.19367cs.AIcs.CV2026-02

探索时间序列与视觉语言在对比表示空间中的对齐极限

Time Series, Vision, and Language: Exploring the Limits of Alignment in Contrastive Representation Spaces

  • 用对比学习对齐独立预训练的时序、视觉和语言编码器
  • 模型越大对齐越好,但时序更易对齐视觉而非文本
  • 丰富描述提升对齐有限,密度超过阈值不再有效

柏拉图表征假说认为,不同模态模型学习到的表征会收敛到世界的共享潜在结构。然而该假说主要在视觉和语言领域验证,时序数据是否参与此收敛尚不明确。我们首先在三模态设置下发现,未显式耦合的时序、视觉和语言编码器在表示空间中呈现近乎正交的几何关系。随后通过在冻结编码器上训练投影头进行后处理对齐,分析其几何结构、缩放行为以及对信息密度和模态特性的依赖。结果表明:整体对齐随模型规模提升,但存在不对称性——时序更易对齐视觉而非文本,图像可作为时序与语言之间的有效中介。此外,更丰富的文本描述仅在达到阈值前促进对齐,密集标注训练不再带来提升;视觉表征亦呈现类似现象。研究为构建包含非传统模态的多模态系统提供了重要启示。

原文摘要 · Abstract (English)

The Platonic Representation Hypothesis posits that learned representations from models trained on different modalities converge to a shared latent structure of the world. However, this hypothesis has largely been examined in vision and language, and it remains unclear whether time series participate in such convergence. We first examine this in a trimodal setting and find that independently pretrained time series, vision, and language encoders exhibit near-orthogonal geometry in the absence of explicit coupling. We then apply post-hoc alignment by training projection heads over frozen encoders using contrastive learning, and analyze the resulting representations with respect to geometry, scaling behavior, and dependence on information density and input modality characteristics. Our investigation reveals that overall alignment in contrastive representation spaces improves with model size, but this alignment is asymmetric: time series align more strongly with visual representations than with text, and images can act as effective intermediaries between time series and language. We further see that richer textual descriptions improve alignment only up to a threshold; training on denser captions does not lead to further improvement. Analogous effects are observed for visual representations. Our findings shed light on considerations for building multimodal systems involving non-conventional data modalities beyond vision and language.

多模态时序数据对比学习表征对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。