arXiv:2605.29462cs.CVcs.AI2026-05ACL

构建中文金融多模态评测集,全面测试大模型在真实场景中的理解能力。

Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset

论文配图:Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset
图 1 · 摘自论文原文
  • 设计覆盖8类金融图像的多模态评测基准CFMME。
  • 顶尖模型问答任务准确率仅66.11%,信息提取平均分77.18。
  • 揭示模型跨模态缺陷,助力金融领域大模型改进。

大型视觉语言模型(LVLMs)已显著拓展模型能力,实现视觉与文本的统一推理,支持更广泛的实际应用。为全面评估LVLMs在中文金融业务全流程中的感知、理解、推理与认知能力,我们提出了全新的中文金融多模态评测基准CFMME。CFMME包含6,052个实例,涵盖从基础学术知识到复杂实际应用,涉及八种主要金融图像模态和四种核心多模态任务。我们在CFMME上对代表性LVLMs进行了全面评估,结果显示,当前最优模型在问答任务上的整体准确率为66.11%,在检测、识别与信息抽取任务上的平均得分为77.18,表明现有模型仍有巨大提升空间。此外,我们还深入分析了错误原因、跨模态能力及多方向设置,为未来研究提供了宝贵洞见。我们期望CFMME能推动LVLMs在金融领域多任务表现的持续进步。

原文摘要 · Abstract (English)

The emergence of Large Vision-Language Models (LVLMs) has substantially expanded model capabilities beyond text-only understanding, enabling unified inference across both visual and textual modalities and supporting a broader range of real-world applications. To comprehensively evaluate the perception, understanding, reasoning, and cognition capabilities of LVLMs throughout the entire financial business workflow in Chinese contexts, we introduce CFMME, a novel Chinese financial multimodal evaluation benchmark. CFMME comprises 6,052 instances spanning from fundamental academic knowledge to complex real-world applications, covering eight primary financial image modalities and four core multimodal tasks. On CFMME, we conduct a thorough evaluation of representative LVLMs. The results show that the state-of-the-art model attains an overall accuracy of 66.11\% on the question answering task and an average score of 77.18 on the detection, recognition, and information extraction tasks, indicating substantial room for improvement in current LVLMs. In addition, we conduct detailed analyses of error causes, cross-modal capabilities, and multi-orientation settings, yielding valuable insights for future research. We hope that CFMME will spur further progress in LVLMs, especially by improving their performance on multiple multimodal tasks in the financial domain.

多模态金融AI评测基准视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。