arXiv:2508.04625cs.CVcs.CE2025-08ICCV被引 10

构建金融多模态推理新基准,提升模型对图表与文本的精准分析能力。

FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging

  • 融合文本与复杂金融图表,设计4.3千道多模态题目
  • 涵盖14个金融领域,难题准确率仅53.0%
  • 适合评估金融场景下多模态模型的深度推理能力

我们提出FinMMR,一个面向金融数值推理的新型双语多模态基准,用于评估多模态大语言模型(MLLMs)在金融领域的推理能力。相比现有基准,本工作实现三项关键改进:(1) 多模态性:系统化转化现有金融推理数据集,并基于最新中文金融研究报告构建新题,包含4.3K问题与8.7K图像,覆盖表格、柱状图、股权结构图等14类;(2) 全面性:涵盖公司金融、银行、行业分析等14个子领域,显著扩展金融知识广度;(3) 挑战性:要求模型结合金融知识,对复杂图文进行多步精确数值推理。当前最佳MLLM在困难题上准确率仅为53.0%。我们认为FinMMR将推动多模态模型在真实金融场景下的推理能力提升。

原文摘要 · Abstract (English)

We present FinMMR, a novel bilingual multimodal benchmark tailored to evaluate the reasoning capabilities of multimodal large language models (MLLMs) in financial numerical reasoning tasks. Compared to existing benchmarks, our work introduces three significant advancements. (1) Multimodality: We meticulously transform existing financial reasoning benchmarks, and construct novel questions from the latest Chinese financial research reports. FinMMR comprises 4.3K questions and 8.7K images spanning 14 categories, including tables, bar charts, and ownership structure charts. (2) Comprehensiveness: FinMMR encompasses 14 financial subdomains, including corporate finance, banking, and industry analysis, significantly exceeding existing benchmarks in financial domain knowledge breadth. (3) Challenge: Models are required to perform multi-step precise numerical reasoning by integrating financial knowledge with the understanding of complex financial images and text. The best-performing MLLM achieves only 53.0% accuracy on Hard problems. We believe that FinMMR will drive advancements in enhancing the reasoning capabilities of MLLMs in real-world scenarios.

金融推理多模态大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。