arXiv:2412.12567cs.CL2024-12ACL被引 11

评测金融领域跨模态多跳推理能力,发现顶尖模型准确率仅30.4%

FCMR: Robust Evaluation of Financial Cross-Modal Multi-Hop Reasoning

  • 构建金融跨模态多跳推理基准FCMR,分三级难度
  • 最难任务中顶尖模型准确率仅30.4%,需三跳跨模态推理
  • 揭示信息检索环节是模型关键瓶颈,适合金融与多模态研究者

现实决策常需整合多模态信息进行推理。尽管近年多模态大语言模型(MLLMs)在该任务上展现出潜力,其跨多种数据源的多跳推理能力仍缺乏充分评估。现有基准如MMQA存在数据污染和复杂查询不足的问题,难以准确衡量性能。为此,我们提出金融跨模态多跳推理(FCMR)基准,旨在通过文本报告、表格和图表等多模态数据,评估MLLMs的推理能力。FCMR分为易、中、难三个层级,其中难题要求精确的跨模态三跳推理,防止忽略任一模态。实验表明,即使最先进模型(Claude 3.5 Sonnet)在最难点也仅达30.4%准确率。我们进一步分析模型内部机制,发现信息检索阶段存在关键瓶颈。

原文摘要 · Abstract (English)

Real-world decision-making often requires integrating and reasoning over information from multiple modalities. While recent multimodal large language models (MLLMs) have shown promise in such tasks, their ability to perform multi-hop reasoning across diverse sources remains insufficiently evaluated. Existing benchmarks, such as MMQA, face challenges due to (1) data contamination and (2) a lack of complex queries that necessitate operations across more than two modalities, hindering accurate performance assessment. To address this, we present Financial Cross-Modal Multi-Hop Reasoning (FCMR), a benchmark created to analyze the reasoning capabilities of MLLMs by urging them to combine information from textual reports, tables, and charts within the financial domain. FCMR is categorized into three difficulty levels-Easy, Medium, and Hard-facilitating a step-by-step evaluation. In particular, problems at the Hard level require precise cross-modal three-hop reasoning and are designed to prevent the disregard of any modality. Experiments on this new benchmark reveal that even state-of-the-art MLLMs struggle, with the best-performing model (Claude 3.5 Sonnet) achieving only 30.4% accuracy on the most challenging tier. We also conduct analysis to provide insights into the inner workings of the models, including the discovery of a critical bottleneck in the information retrieval phase.

多模态金融推理评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。