构建多轮多模态金融推理数据集,评测模型真实场景下的分析能力
FinMTM: A Multi-Turn Multimodal Benchmark for Financial Reasoning and Agent Evaluation
- 基于中英文金融图表构建1.1万组多轮问答对
- 22个视觉语言模型在长上下文推理上表现不佳
- 专为智能投顾等复杂任务设计评估方案
金融领域对视觉语言模型(VLMs)构成严峻挑战,源于专业图表格式与知识密集型推理需求。现有金融基准大多为单轮且题型单一,难以全面评估真实应用场景。为此,我们提出FinMTM,一个在数据和任务维度均拓展的多轮多模态基准。数据方面,我们整理并标注了11,133组双语(中文和英文)金融问答对,基于蜡烛图、统计图和报告图表等金融视觉内容。任务方面,涵盖单选、多选、多轮开放式对话及基于代理的任务。我们还设计了特定任务评估协议:多选题采用集合重叠评分规则,多轮对话使用逐轮与会话级得分加权组合,代理任务则综合规划质量与最终结果制定复合指标。对22个VLMs的广泛实验揭示其在细粒度视觉感知、长上下文推理和复杂代理流程中的局限性。
原文摘要 · Abstract (English)
The financial domain poses substantial challenges for vision-language models (VLMs) due to specialized chart formats and knowledge-intensive reasoning requirements. However, existing financial benchmarks are largely single-turn and rely on a narrow set of question formats, limiting comprehensive evaluation in realistic application scenarios. To address this gap, we propose FinMTM, a multi-turn multimodal benchmark that expands diversity along both data and task dimensions. On the data side, we curate and annotate 11{,}133 bilingual (Chinese and English) financial QA pairs grounded in financial visuals, including candlestick charts, statistical plots, and report figures. On the task side, FinMTM covers single- and multiple-choice questions, multi-turn open-ended dialogues, and agent-based tasks. We further design task-specific evaluation protocols, including a set-overlap scoring rule for multiple-choice questions, a weighted combination of turn-level and session-level scores for multi-turn dialogues, and a composite metric that integrates planning quality with final outcomes for agent tasks. Extensive experimental evaluation of 22 VLMs reveal their limitations in fine-grained visual perception, long-context reasoning, and complex agent workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。