构建中文多模态金融评测集,评估大模型处理图表与文本的财务分析能力。
CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model
- 设计包含9000+图文对的中文金融多模态数据集,覆盖多种图表类型。
- 提出分阶段评测机制,逐步增加视觉内容复杂度以精准评估模型表现。
- 发现模型普遍存在图表误读和金融概念理解错误,需针对性优化。
多模态大语言模型(MLLMs)随着大语言模型的发展迅速演进,并已应用于多个领域。在金融领域,文本、图表和表格等多模态信息的融合对准确高效决策至关重要。因此,建立一个涵盖多种数据类型的评测体系对推动金融应用具有重要意义。本文提出CFBenchmark-MM,一个包含超过9,000个图像-问题对的中文多模态金融基准数据集,涵盖表格、柱状图、折线图、饼图及结构图等多种可视化形式。此外,我们设计了分阶段评测机制,通过逐步引入视觉内容来评估模型处理多模态信息的能力。尽管大模型具备一定的金融知识,实验结果仍显示其在处理多模态金融语境时效率与鲁棒性有限。进一步分析表明,错误主要源于对视觉内容的误读和对金融概念的理解偏差。本研究验证了多模态大模型在金融分析中巨大的潜在价值,但目前尚未被充分开发,亟需进一步发展与领域特化优化。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have rapidly evolved with the growth of Large Language Models (LLMs) and are now applied in various fields. In finance, the integration of diverse modalities such as text, charts, and tables is crucial for accurate and efficient decision-making. Therefore, an effective evaluation system that incorporates these data types is essential for advancing financial application. In this paper, we introduce CFBenchmark-MM, a Chinese multimodal financial benchmark with over 9,000 image-question pairs featuring tables, histogram charts, line charts, pie charts, and structural diagrams. Additionally, we develop a staged evaluation system to assess MLLMs in handling multimodal information by providing different visual content step by step. Despite MLLMs having inherent financial knowledge, experimental results still show limited efficiency and robustness in handling multimodal financial context. Further analysis on incorrect responses reveals the misinterpretation of visual content and the misunderstanding of financial concepts are the primary issues. Our research validates the significant, yet underexploited, potential of MLLMs in financial analysis, highlighting the need for further development and domain-specific optimization to encourage the enhanced use in financial domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。