首次系统评估视觉大模型在时序分析中的表现,发现分类有效但预测受限。
From Images to Signals: Are Large Vision Models Useful for Time Series Analysis?
- 用图像化方法将时序数据转为图像,输入四大视觉大模型测试
- 分类任务表现优异,但预测任务受制于特定模型与窗口长度
- 揭示视觉大模型在长周期预测中存在偏差和能力局限,适合研究多模态时序建模者
基于Transformer的模型在时序研究中日益受到关注,推动了大语言模型(LLMs)和基础模型在时序分析中的应用。随着多模态趋势发展,大型视觉模型(LVMs)正成为有前景的方向。过去,Transformer和LLMs在时序领域的有效性存在争议。对于LVMs,同样存在疑问:它们是否真正适用于时序分析?为此,我们设计并开展首个系统性研究,涵盖4个LVM、8种图像化方法、18个数据集和26个基线,在高阶(分类)与低阶(预测)任务上进行评估,并包含详尽的消融分析。结果表明,LVMs确实对时序分类有用,但在预测任务中面临挑战。尽管有效,当前最优的LVM预测模型仅适用于特定类型的模型和图像化方法,对预测周期存在偏差,且难以利用长回顾窗口。期望本研究能为未来基于LVM和多模态的时序分析研究奠定基础。
原文摘要 · Abstract (English)
Transformer-based models have gained increasing attention in time series research, driving interest in Large Language Models (LLMs) and foundation models for time series analysis. As the field moves toward multi-modality, Large Vision Models (LVMs) are emerging as a promising direction. In the past, the effectiveness of Transformer and LLMs in time series has been debated. When it comes to LVMs, a similar question arises: are LVMs truely useful for time series analysis? To address it, we design and conduct the first principled study involving 4 LVMs, 8 imaging methods, 18 datasets and 26 baselines across both high-level (classification) and low-level (forecasting) tasks, with extensive ablation analysis. Our findings indicate LVMs are indeed useful for time series classification but face challenges in forecasting. Although effective, the contemporary best LVM forecasters are limited to specific types of LVMs and imaging methods, exhibit a bias toward forecasting periods, and have limited ability to utilize long look-back windows. We hope our findings could serve as a cornerstone for future research on LVM- and multimodal-based solutions to different time series tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。