构建金融多模态推理新基准,挑战模型跨文档多步计算与场景理解能力。
FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation
- 设计包含12类隐含金融场景的1200道题,逼迫模型做专家级推理。
- 引入9类共837份中英文长文档,平均50.8页,含丰富视觉信息。
- 平均需11步推理,65%问题跨页找证据,当前最佳模型仅58%准确率。
我们提出FinMMDocR,一个用于评估多模态大语言模型(MLLMs)在真实金融数值推理任务上的新型双语多模态基准。相比现有基准,本工作实现三大突破:(1) 场景意识:1200道专家标注题中57.9%包含12类隐含金融场景(如投资组合管理),要求模型基于假设进行专家级推理;(2) 文档理解:涵盖9类中英文文档共837份,平均50.8页,包含丰富视觉元素,在文档广度与深度上显著超越现有基准;(3) 多步计算:题目平均需11步推理(5.3步提取 + 5.7步计算),65.0%需跨页证据(平均2.4页)。目前最优的MLLM仅达58.0%准确率,不同检索增强生成(RAG)方法表现差异显著。我们期望FinMMDocR能推动MLLM及推理增强方法在真实复杂多模态任务上的进步。
原文摘要 · Abstract (English)
We introduce FinMMDocR, a novel bilingual multimodal benchmark for evaluating multimodal large language models (MLLMs) on real-world financial numerical reasoning. Compared to existing benchmarks, our work delivers three major advancements. (1) Scenario Awareness: 57.9% of 1,200 expert-annotated problems incorporate 12 types of implicit financial scenarios (e.g., Portfolio Management), challenging models to perform expert-level reasoning based on assumptions; (2) Document Understanding: 837 Chinese/English documents spanning 9 types (e.g., Company Research) average 50.8 pages with rich visual elements, significantly surpassing existing benchmarks in both breadth and depth of financial documents; (3) Multi-Step Computation: Problems demand 11-step reasoning on average (5.3 extraction + 5.7 calculation steps), with 65.0% requiring cross-page evidence (2.4 pages average). The best-performing MLLM achieves only 58.0% accuracy, and different retrieval-augmented generation (RAG) methods show significant performance variations on this task. We expect FinMMDocR to drive improvements in MLLMs and reasoning-enhanced methods on complex multimodal reasoning tasks in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。