构建金融多模态推理基准,提升AI理解图表与文本的能力。
Understanding Financial Reasoning in AI: A Multimodal Benchmark and Error Learning Approach
- 设计包含3200个问题的金融多模态评测集,融合文本与视觉数据。
- 引入错误反馈机制,无需微调即可提升模型推理准确率。
- 揭示当前AI在图表理解与数学逻辑上的不足,适合金融AI研究者参考。
有效的金融推理不仅需要文本理解能力,还需解析复杂视觉信息,如图表、表格和趋势图。本文提出一个新基准,用于评估大语言模型与多模态模型在金融场景下的推理表现。该基准涵盖15个核心金融主题的3,200个专家级问答对,整合文本与视觉模态,真实反映金融分析中的挑战。为克服现有推理方法局限,我们提出一种基于历史错误与反馈的自省学习框架,无需微调即可引导推理过程。实验表明,多模态输入显著提升性能,且引入错误反馈可带来持续且可量化的改进。结果揭示了模型在视觉理解与数学逻辑方面的长期瓶颈,同时展示了自反思推理在金融AI中的潜力。代码与数据详见 https://anonymous/FinMR/CodeData。
原文摘要 · Abstract (English)
Effective financial reasoning demands not only textual understanding but also the ability to interpret complex visual data such as charts, tables, and trend graphs. This paper introduces a new benchmark designed to evaluate how well AI models - especially large language and multimodal models - reason in finance-specific contexts. Covering 3,200 expert-level question-answer pairs across 15 core financial topics, the benchmark integrates both textual and visual modalities to reflect authentic analytical challenges in finance. To address limitations in current reasoning approaches, we propose an error-aware learning framework that leverages historical model mistakes and feedback to guide inference, without requiring fine-tuning. Our experiments across state-of-the-art models show that multimodal inputs significantly enhance performance and that incorporating error feedback leads to consistent and measurable improvements. The results highlight persistent challenges in visual understanding and mathematical logic, while also demonstrating the promise of self-reflective reasoning in financial AI systems. Our code and data can be found at https://anonymous/FinMR/CodeData.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。