构建金融领域专家级多模态理解基准,评测模型对图表与专业知识的推理能力。
MME-Finance: A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning
- 针对金融场景设计真实用户使用图像与专业问题,由10年以上从业者标注。
- 19个主流多模态大模型在该基准上表现不佳,顶尖模型仅达65.69分。
- 专为中文金融场景优化,支持中英文对比评估,适合金融AI研究者使用。
近年来,通用领域的多模态基准推动了多模态模型在通用任务上的快速发展。然而,金融领域具有独特性:包含独特的图形(如蜡烛图、技术指标图)和丰富的专业知识(如期货、换手率)。通用基准难以衡量多模态模型在金融领域的性能,无法有效引导大型金融模型的发展。为促进大型金融多模态模型的发展,我们提出 MME-Finance,一个双语、开放、面向实际应用的视觉问答(VQA)基准。其特点在于金融性与专业性:构建反映真实用户需求的图表(如电脑截图、手机拍摄),根据金融领域提问偏好设计问题,并由具有10年以上金融行业经验的专家进行标注。此外,我们设计了定制化的金融评估体系,将视觉信息引入多模态评估流程。对19个主流多模态大语言模型进行了广泛实验,测试其感知、推理与认知能力。结果表明,通用基准表现优异的模型在MME-Finance上表现不佳;例如,开源与闭源模型最高分分别为65.69(Qwen2VL-72B)和63.18(GPT-4o)。尤其在与金融密切相关的蜡烛图和技术指标图类别上表现较差。同时,我们提出了中文版本,有助于在中文语境下比较多模态大模型的性能。
原文摘要 · Abstract (English)
In recent years, multimodal benchmarks for general domains have guided the rapid development of multimodal models on general tasks. However, the financial field has its peculiarities. It features unique graphical images (e.g., candlestick charts, technical indicator charts) and possesses a wealth of specialized financial knowledge (e.g., futures, turnover rate). Therefore, benchmarks from general fields often fail to measure the performance of multimodal models in the financial domain, and thus cannot effectively guide the rapid development of large financial models. To promote the development of large financial multimodal models, we propose MME-Finance, an bilingual open-ended and practical usage-oriented Visual Question Answering (VQA) benchmark. The characteristics of our benchmark are finance and expertise, which include constructing charts that reflect the actual usage needs of users (e.g., computer screenshots and mobile photography), creating questions according to the preferences in financial domain inquiries, and annotating questions by experts with 10+ years of experience in the financial industry. Additionally, we have developed a custom-designed financial evaluation system in which visual information is first introduced in the multi-modal evaluation process. Extensive experimental evaluations of 19 mainstream MLLMs are conducted to test their perception, reasoning, and cognition capabilities. The results indicate that models performing well on general benchmarks cannot do well on MME-Finance; for instance, the top-performing open-source and closed-source models obtain 65.69 (Qwen2VL-72B) and 63.18 (GPT-4o), respectively. Their performance is particularly poor in categories most relevant to finance, such as candlestick charts and technical indicator charts. In addition, we propose a Chinese version, which helps compare performance of MLLMs under a Chinese context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。