构建专业级金融多模态推理数据集,评测模型真实分析能力
FinMR: A Knowledge-Intensive Multimodal Benchmark for Advanced Financial Reasoning
- 构建含3200+题对的多模态金融数据集,覆盖15类专业主题
- 模型在复杂公式应用与图像解析上显著落后于人类分析师
- 适合评估金融智能系统、开发高阶金融大模型的研究者使用
近年来,多模态大语言模型取得了显著进展,但在金融等专业领域中,其严格评估仍受限于缺乏具备专业级知识强度、详尽标注和高级推理复杂性的数据集。为填补这一关键空白,我们提出FinMR,一个高质量、知识密集型的多模态数据集,专为评估专业分析师水平的金融推理能力而设计。FinMR包含超过3,200个精心筛选并由专家标注的问答对,覆盖15种多样化的金融主题,确保领域多样性,并融合了复杂的数学推理、高级金融知识及多种图像类型的细微视觉理解任务。通过在领先闭源与开源多模态大模型上进行全面基准测试,我们揭示了这些模型与专业金融分析师之间显著的性能差距,暴露了模型在精确图像分析、复杂金融公式准确应用以及深层上下文理解等方面的不足。凭借丰富多样的视觉内容和详尽的解释性标注,FinMR成为评估和推动多模态金融推理向专业分析师水平迈进的关键基准工具。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have made substantial progress in recent years. However, their rigorous evaluation within specialized domains like finance is hindered by the absence of datasets characterized by professional-level knowledge intensity, detailed annotations, and advanced reasoning complexity. To address this critical gap, we introduce FinMR, a high-quality, knowledge-intensive multimodal dataset explicitly designed to evaluate expert-level financial reasoning capabilities at a professional analyst's standard. FinMR comprises over 3,200 meticulously curated and expertly annotated question-answer pairs across 15 diverse financial topics, ensuring broad domain diversity and integrating sophisticated mathematical reasoning, advanced financial knowledge, and nuanced visual interpretation tasks across multiple image types. Through comprehensive benchmarking with leading closed-source and open-source MLLMs, we highlight significant performance disparities between these models and professional financial analysts, uncovering key areas for model advancement, such as precise image analysis, accurate application of complex financial formulas, and deeper contextual financial understanding. By providing richly varied visual content and thorough explanatory annotations, FinMR establishes itself as an essential benchmark tool for assessing and advancing multimodal financial reasoning toward professional analyst-level competence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。