arXiv:2411.06018cs.LGcs.AI2024-11NAACL被引 43

让大模型通过看图理解时间序列,性能提升140%。

A Picture is Worth A Thousand Numbers: Enabling LLMs Reason about Time Series via Visualization

  • 用可视化图像辅助大模型进行时间序列推理
  • 零样本推理性能提升140%,令牌消耗降低99%
  • 适合需要高效时序分析的AI研究者和工程师

大型语言模型(LLMs)在多个领域展现出强大的推理能力,但在时间序列推理(TsR)方面仍鲜有探索,而时序数据在现实世界中普遍存在。本文提出TimerBed,首个全面评估LLMs时序推理能力的测试基准。该基准包含分层的推理模式、真实世界任务、多种LLM与推理策略组合,以及监督模型作为对比基准。通过大量实验,验证了LLMs在时序推理中的初始失败:零样本(ZST)无效,少量样本上下文学习(ICL)性能下降。我们识别出可能的根本原因:数据的数值建模问题。为此,提出基于提示的解决方案VL-Time,利用可视化建模数据结合语言引导推理。实验表明,VL-Time使多模态LLMs成为非平凡的零样本推理者和强大的上下文学习者,平均性能提升约140%,平均令牌消耗减少99%。

原文摘要 · Abstract (English)

Large language models (LLMs), with demonstrated reasoning abilities across multiple domains, are largely underexplored for time-series reasoning (TsR), which is ubiquitous in the real world. In this work, we propose TimerBed, the first comprehensive testbed for evaluating LLMs' TsR performance. Specifically, TimerBed includes stratified reasoning patterns with real-world tasks, comprehensive combinations of LLMs and reasoning strategies, and various supervised models as comparison anchors. We perform extensive experiments with TimerBed, test multiple current beliefs, and verify the initial failures of LLMs in TsR, evidenced by the ineffectiveness of zero shot (ZST) and performance degradation of few shot in-context learning (ICL). Further, we identify one possible root cause: the numerical modeling of data. To address this, we propose a prompt-based solution VL-Time, using visualization-modeled data and language-guided reasoning. Experimental results demonstrate that Vl-Time enables multimodal LLMs to be non-trivial ZST and powerful ICL reasoners for time series, achieving about 140% average performance improvement and 99% average token costs reduction.

时间序列视觉推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。