arXiv:2505.24714cs.CL2025-05ACL被引 44

金融多模态评测新基准,涵盖超万份高质量研报数据。

FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation

  • 构建11000+样本的金融多模态数据集,覆盖18个领域与6类资产。
  • 引入防幻觉评分机制,模型在该基准上表现仍不理想。
  • 适合金融AI、多模态大模型研究者使用,评测可靠性强。

近年来多模态大语言模型(MLLMs)发展迅速,但在金融领域缺乏有效的专用多模态评估数据集。为推动金融领域MLLMs的发展,我们提出FinMME,包含超过11,000个高质量金融研究样本,覆盖18个金融领域和6种资产类别,包含10类主要图表类型及21种子类型。通过20名标注员和精心设计的验证机制保障数据质量。同时开发了包含幻觉惩罚与多维度能力评估的FinScore评估体系,实现无偏评测。大量实验表明,即使最先进的模型如GPT-4o在FinMME上表现也欠佳,凸显其挑战性。该基准在不同提示下预测波动低于1%,展现出优异的鲁棒性。数据集与评估协议已开源:https://huggingface.co/datasets/luojunyu/FinMME 与 https://github.com/luo-junyu/FinMME。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have experienced rapid development in recent years. However, in the financial domain, there is a notable lack of effective and specialized multimodal evaluation datasets. To advance the development of MLLMs in the finance domain, we introduce FinMME, encompassing more than 11,000 high-quality financial research samples across 18 financial domains and 6 asset classes, featuring 10 major chart types and 21 subtypes. We ensure data quality through 20 annotators and carefully designed validation mechanisms. Additionally, we develop FinScore, an evaluation system incorporating hallucination penalties and multi-dimensional capability assessment to provide an unbiased evaluation. Extensive experimental results demonstrate that even state-of-the-art models like GPT-4o exhibit unsatisfactory performance on FinMME, highlighting its challenging nature. The benchmark exhibits high robustness with prediction variations under different prompts remaining below 1%, demonstrating superior reliability compared to existing datasets. Our dataset and evaluation protocol are available at https://huggingface.co/datasets/luojunyu/FinMME and https://github.com/luo-junyu/FinMME.

金融AI多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。