用视觉语言模型评估时间序列预测,更贴近人类直觉。
TimeVista: Exploring and Exploiting Vision-Language Models as Judges for Time Series Forecasting

- 用视觉语言模型分析时序图与文本描述,结合上下文判断预测质量。
- 在5563个样本上验证,模型评分与人类偏好一致性显著高于传统指标。
- 适合关注预测结果可解释性与人类感知的时序建模研究者。
高质量的时间序列预测对实际决策至关重要。然而,传统的点对点评估指标往往无法揭示复杂的时序模式,且与人类直观偏好不符。尽管大语言模型作为评判者(LLM-as-a-Judge)已革新文本评价,但其在时间序列领域的应用仍基本空白。本文提出利用视觉语言模型(VLMs)作为时间序列预测的评判者,借助其理解时序图表与文本信息的能力。我们构建了包含5563个时间序列样本及详细评估标准的TimeVista基准,融合微观与宏观层面的上下文判断。大量元评估表明,VLMs作为评判者具有高度可靠性,其评分与人类偏好一致性显著优于传统指标。基于该基准,我们系统评估了近期时间序列基础模型(TSFMs),结果表明VLMs能提供稳健、可解释且与人类对齐的评估标准。
原文摘要 · Abstract (English)
High-quality time series forecasting is pivotal for real-world decision-making. However, traditional point-wise metrics often fail to reveal complex temporal patterns and align poorly with human intuitive preferences. While the ''LLM-as-a-Judge'' paradigm has revolutionized text evaluation by providing flexible, human-aligned judgment, its application to time series remains largely unexplored. In this paper, we leverage Vision-Language Models (VLMs) as judges for time series forecasting, harnessing their ability to comprehend time series plots grounded in textual information. Specifically, we propose a novel framework integrating micro- and macro-level judgments informed by contextual information to evaluate time series forecasting. To this end, we introduce TimeVista, a comprehensive VLM-as-a-Judge benchmark comprising 5563 time series samples paired with detailed evaluation rubrics. Extensive meta-evaluations demonstrate that VLMs are highly reliable judges, achieving significantly higher consistency with human preferences than conventional metrics. Building upon our benchmark, we comprehensively assess recent Time Series Foundation Models (TSFMs) under the VLM-as-a-Judge paradigm. Our results demonstrate that VLMs serve as robust and interpretable judges, providing a comprehensive, human-aligned standard for evaluating time series models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。