arXiv:2410.04526cs.CLcs.AI2024-10被引 22

首个金融领域多语言多模态问答基准,评测大模型复杂推理能力。

FAMMA: A Benchmark for Financial Domain Multilingual Multimodal Question Answering

  • 构建双版本金融多模态问答数据集,含1945题基础题与103题专家新题。
  • 覆盖8个金融子领域,支持中/英/法三语,每题配图表等非文本信息。
  • 揭示当前大模型在金融复杂推理上仍有显著不足,提示可借助推理轨迹提升性能。

本文提出FAMMA,一个开源的金融领域多语言多模态问答基准。该基准旨在评估大语言模型(LLMs)在回答需要高级金融知识的复杂推理问题时的能力。基准包含两个版本:FAMMA-Basic 包含从大学教材和考试中提取的1,945道题目,附有人工标注的答案与推理过程;FAMMA-LivePro 包含由领域专家创作的103道新题,答案与推理过程暂不公开,确保评估无污染。题目覆盖金融学八大核心子领域(如公司金融、衍生品、投资组合管理),部分为中文或法语,多数为英语。每道题均包含图表、流程图或表格等非文本数据。实验表明,FAMMA对包括GPT-o1和DeepSeek-R1在内的主流推理模型构成显著挑战。我们还收集了DeepSeek-R1在FAMMA-Basic上的1,270条推理轨迹,并使用这些数据微调了一系列开源Qwen模型。结果显示,基于推理轨迹训练可显著提升模型在FAMMA-LivePro上的表现。相关榜单、数据、代码及训练模型已发布于 https://famma-bench.github.io/famma/。

原文摘要 · Abstract (English)

In this paper, we introduce FAMMA, an open-source benchmark for \underline{f}in\underline{a}ncial \underline{m}ultilingual \underline{m}ultimodal question \underline{a}nswering (QA). Our benchmark aims to evaluate the abilities of large language models (LLMs) in answering complex reasoning questions that require advanced financial knowledge. The benchmark has two versions: FAMMA-Basic consists of 1,945 questions extracted from university textbooks and exams, along with human-annotated answers and rationales; FAMMA-LivePro consists of 103 novel questions created by human domain experts, with answers and rationales held out from the public for a contamination-free evaluation. These questions cover advanced knowledge of 8 major subfields in finance (e.g., corporate finance, derivatives, and portfolio management). Some are in Chinese or French, while a majority of them are in English. Each question has some non-text data such as charts, diagrams, or tables. Our experiments reveal that FAMMA poses a significant challenge on LLMs, including reasoning models such as GPT-o1 and DeepSeek-R1. Additionally, we curated 1,270 reasoning trajectories of DeepSeek-R1 on the FAMMA-Basic data, and fine-tuned a series of open-source Qwen models using this reasoning data. We found that training a model on these reasoning trajectories can significantly improve its performance on FAMMA-LivePro. We released our leaderboard, data, code, and trained models at https://famma-bench.github.io/famma/.

金融AI多模态大模型评测多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。