arXiv:2507.14823cs.CV2025-07

首个专用于金融图表理解的评测基准,揭示视觉语言模型的短板

FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models

  • 构建1200张真实金融图表数据集,含7016个问答任务
  • 开源与闭源模型差距缩小,但空间推理能力仍弱
  • 适合研究金融视觉理解或模型评估的学者参考

大型视觉语言模型在图表理解方面取得显著进展,但具有复杂时间结构和领域术语的金融图表仍鲜少被深入研究。本文提出FinChart-Bench,首个专注于真实金融图表的基准测试。该数据集包含2015至2024年间收集的1,200张金融图表图像,每张图配有真/假(TF)、多选(MC)和问答(QA)类问题,共7,016个问题。我们对25个顶尖视觉语言模型进行了全面评估。结果揭示:(1)开源与闭源模型性能差距正在缩小;(2)同系列模型升级后表现反而下降;(3)多数模型指令遵循能力不足;(4)先进模型在空间推理上仍有明显局限;(5)当前模型尚不可靠作为自动化评估工具。这些发现凸显了现有视觉语言模型在金融图表理解上的关键瓶颈。数据集已公开于https://huggingface.co/datasets/Tizzzzy/FinChart-Bench。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) have made significant progress in chart understanding. However, financial charts, characterized by complex temporal structures and domain-specific terminology, remain notably underexplored. We introduce FinChart-Bench, the first benchmark specifically focused on real-world financial charts. FinChart-Bench comprises 1,200 financial chart images collected from 2015 to 2024, each annotated with True/False (TF), Multiple Choice (MC), and Question Answering (QA) questions, totaling 7,016 questions. We conduct a comprehensive evaluation of 25 state-of-the-art LVLMs on FinChart-Bench. Our evaluation reveals critical insights: (1) the performance gap between open-source and closed-source models is narrowing, (2) performance degradation occurs in upgraded models within families, (3) many models struggle with instruction following, (4) both advanced models show significant limitations in spatial reasoning abilities, and (5) current LVLMs are not reliable enough to serve as automated evaluators. These findings highlight important limitations in current LVLM capabilities for financial chart understanding. The FinChart-Bench dataset is available at https://huggingface.co/datasets/Tizzzzy/FinChart-Bench.

金融图表视觉语言模型评测基准空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。