arXiv:2506.13066cs.CL2025-06被引 3

构建金融多模态数据集并优化奖励机制,提升大模型推理能力。

FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design

  • 自动化流水线生成8.9万组图文对,解决财务报告图文错位问题。
  • 引入对抗性奖励与多图像对比学习,提升模型视觉理解与逻辑推理。
  • 在7个基准上显著超越现有模型,适合金融智能分析场景使用。

大型多模态模型(LMMs)展现出强大的跨模态推理能力,但金融应用受限于高质量多模态推理数据稀缺及训练范式效率低下。为此,我们提出集成框架FinLMM-R1,结合自动化的可扩展数据构建流水线与增强型训练策略,以提升LMM的多模态推理性能。自动化与可扩展流水线(ASP)通过独立的问答生成与图文对齐机制,解决财务报告中的文本-视觉错位问题,确保数据完整性与提取效率。基于ASP,我们从23,397份财务报告中收集了89,378组对齐的图像-问题对,涵盖算术推理、统计推理、财务解释与财务知识等任务。同时,我们提出思考过程对抗性奖励框架(TAR-LMM),在两阶段训练基础上增加奖励机制:第一阶段聚焦纯文本任务,采用格式与准确率奖励引导模型生成结构化推理内容;第二阶段构建多图像对比样本,引入图像选择、思维长度与对抗性奖励,联合优化模型在视觉感知、推理效率与逻辑连贯性方面的表现。在7个基准上的大量实验表明,基于ASP的数据集与训练框架在通用与金融多模态场景下均显著提升答案准确率与推理深度。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) demonstrate significant cross-modal reasoning capabilities. However, financial applications face challenges due to the lack of high-quality multimodal reasoning datasets and the inefficiency of existing training paradigms for reasoning enhancement. To address these issues, we propose an integrated framework, FinLMM-R1, combining an automated and scalable pipeline for data construction with enhanced training strategies to improve the multimodal reasoning of LMM. The Automated and Scalable Pipeline (ASP) resolves textual-visual misalignment in financial reports through a separate paradigm of question-answer generation and image-question alignment, ensuring data integrity and extraction efficiency. Through ASP, we collect 89,378 aligned image-question pairs from 23,397 financial reports, covering tasks such as arithmetic reasoning, statistics reasoning, financial explanation, and financial knowledge. Moreover, we introduce the Thinking with Adversarial Reward in LMM (TAR-LMM), extending the prior two-stage training framework [1] with additional reward mechanisms. In the first stage, we focus on text-only tasks with format and accuracy rewards to guide the model in generating well-structured thinking contents. In the second stage, we construct multi-image contrastive samples with additional reward components including image selection, thinking content length, and adversarial reward to jointly optimize the LMM across visual perception, reasoning efficiency, and logical coherence. Extensive experiments on 7 benchmarks show ASP-derived dataset and training framework significantly improve answer accuracy and reasoning depth over existing reasoning LMMs in both general and financial multimodal contexts.

多模态推理金融AI数据构建奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。