用视觉语言模型分析股票走势,不训练也能预测得更好。
VISTA: Vision-Language Inference for Training-Free Stock Time-Series Analysis
- 用图表+文字双模态输入,零样本提示预测股价
- 比传统方法和纯文本模型最高提升89.83%
- 适合金融分析、量化研究等无需训练的场景
股票价格预测是金融分析中复杂且高风险的任务,传统上依赖统计模型,近年也采用语言模型。本文提出VISTA(Vision-Language Inference for Stock Time-series Analysis),一种无需训练的多模态股票预测框架,利用视觉-语言模型(VLM)结合历史股价数值与对应折线图进行未来价格预测。通过精心设计的思维链提示,在零样本设定下融合数值与视觉模态,捕捉单模态方法常忽略的互补模式。在标准基线(如ARIMA和仅文本的LLM提示方法)上的实验表明,VISTA性能最高提升89.83%,验证了多模态推理在股票时间序列分析中的有效性,并展示了无需任务训练的VLM在金融预测中的潜力。
原文摘要 · Abstract (English)
Stock price prediction remains a complex and high-stakes task in financial analysis, traditionally addressed using statistical models or, more recently, language models. In this work, we introduce VISTA (Vision-Language Inference for Stock Time-series Analysis), a novel, training-free framework that leverages Vision-Language Models (VLMs) for multi-modal stock forecasting. VISTA prompts a VLM with both textual representations of historical stock prices and their corresponding line charts to predict future price values. By combining numerical and visual modalities in a zero-shot setting and using carefully designed chain-of-thought prompts, VISTA captures complementary patterns that unimodal approaches often miss. We benchmark VISTA against standard baselines, including ARIMA and text-only LLM-based prompting methods. Experimental results show that VISTA outperforms these baselines by up to 89.83%, demonstrating the effectiveness of multi-modal inference for stock time-series analysis and highlighting the potential of VLMs in financial forecasting tasks without requiring task-specific training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。