测试视觉语言模型对多尺度蜡烛图的理解能力,发现其在复杂市场中表现不佳。
Do VLMs Truly "Read" Candlesticks? A Multi-Scale Benchmark for Visual Stock Price Forecasting

- 构建多尺度蜡烛图数据集与评估框架,检验模型对长短周期信号的整合能力。
- 多数模型仅在单边趋势时表现良好,常见市场场景下预测能力弱。
- 模型对提示中的时间跨度不敏感,存在显著预测偏差,适合研究多尺度视觉推理者关注。
视觉语言模型(VLMs)被越来越多地用于视觉股票价格预测,但现有基准未能充分评估其对蜡烛图中股价信息的真实理解。首先,先前研究未能区分模型性能提升是否源于对视觉输入的真实理解,还是仅依赖表格式输入。其次,多数数据集和评估设置基于单一周期或表格输入,而人类分析师高度依赖多尺度蜡烛图——长期周期捕捉趋势方向,短期周期提供反转点线索。因此,难以系统评估VLM对短长期视觉市场动态的融合能力。为此,我们构建了一个多尺度蜡烛图数据集与标准化评估框架,以检验VLM利用多尺度视觉市场信号的能力。评估结合混淆矩阵诊断与信息系数(IC)时间序列指标,并引入基于特征的XGBoost作为时间基线。使用该数据集,我们对代表性VLM进行了基准测试并分析其利用多尺度股价数据的能力。实验结果表明,大多数VLM仅在持续上涨或下跌条件下表现良好,而在更常见的市场情景中预测能力较弱。同时,我们识别出显著的预测偏差及对提示中明确指定的预测时间跨度敏感性不足,表明其在精确时间推理方面存在固有局限。
原文摘要 · Abstract (English)
Vision-language models(VLMs) are increasingly applied to visual stock price forecasting, yet existing benchmarks inadequately evaluate their understanding of stock price in candlestick charts. First, prior studies fail to isolate VLMs' comprehension of visual inputs genuinely improves predictive performance and whether VLMs truly comprehend candlestick patterns. Further, most existing datasets and evaluation setups are designed around single-period or tabular inputs. However, human analysts strongly rely on multi-scale candlestick charts, where longer-term horizons capture trend direction and shorter-term horizons provide cues for inflection points, making it difficult to systematically assess VLMs' ability to integrate short-term and long-term visual market dynamics. To bridge this gap, we construct a multi-scale candlestick charts dataset and a standardized evaluation framework to assess VLMs' ability to utilize multi-scale visual market signals. Evaluation combines confusion-matrix-based diagnostics with information coefficient(IC) time series metrics and includes XGBoost as a feature-based temporal baseline. Using this dataset, we benchmark representative VLMs and analyze their ability to leverage multi-scale stock price data. Experimental results show that most VLMs perform well only under persistent uptrend or downtrend conditions, while exhibiting weak predictive capability in more common market scenarios. We also identify significant prediction biases and limited sensitivity to explicitly specified forecast horizons in prompts, indicating inherent limitations in precise temporal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。