arXiv:2507.06057cs.AIcs.LG2025-07被引 1

让大模型学会金融推理,用三阶段训练提升专业能力。

FEVO: Financial Knowledge Expansion and Reasoning Evolution for Large Language Models

  • 三步增强:预训练扩知识、微调建逻辑、强化学习融推理。
  • 在5个金融基准上超越更大模型,性能达当前最优。
  • 适合需要精准金融分析的从业者与研究者。

大型语言模型(LLM)在数学和编程等领域的推理能力显著提升,但在金融领域因需大量专业知识,相关研究仍有限。为此,我们提出FEVO(Financial Evolution)多阶段增强框架,通过持续预训练(CPT)扩展金融知识,监督微调(SFT)引入结构化推理模式,强化学习(RL)融合知识与推理。为保障训练效率与质量,我们利用前沿推理模型和规则过滤构建了专用于各训练阶段的高质量数据集FEVO-Train。基于该框架,我们从Qwen2.5-32B训练出FEVO系列模型——C32B、S32B、R32B,并在七个基准上评估其金融与通用能力。结果显示,FEVO-R32B在五个金融基准上达到当前最优表现,优于许多更大模型及专用模型;且显著优于仅通过强化学习训练的FEVO-R32B-0,验证了金融知识扩展与结构化逻辑推理蒸馏的有效性。

原文摘要 · Abstract (English)

Advancements in reasoning for large language models (LLMs) have lead to significant performance improvements for LLMs in various fields such as mathematics and programming. However, research applying these advances to the financial domain, where considerable domain-specific knowledge is necessary to complete tasks, remains limited. To address this gap, we introduce FEVO (Financial Evolution), a multi-stage enhancement framework developed to enhance LLM performance in the financial domain. FEVO systemically enhances LLM performance by using continued pre-training (CPT) to expand financial domain knowledge, supervised fine-tuning (SFT) to instill structured, elaborate reasoning patterns, and reinforcement learning (RL) to further integrate the expanded financial domain knowledge with the learned structured reasoning. To ensure effective and efficient training, we leverage frontier reasoning models and rule-based filtering to curate FEVO-Train, high-quality datasets specifically designed for the different post-training phases. Using our framework, we train the FEVO series of models - C32B, S32B, R32B - from Qwen2.5-32B and evaluate them on seven benchmarks to assess financial and general capabilities, with results showing that FEVO-R32B achieves state-of-the-art performance on five financial benchmarks against much larger models as well as specialist models. More significantly, FEVO-R32B demonstrates markedly better performance than FEVO-R32B-0 (trained from Qwen2.5-32B-Instruct using only RL), thus validating the effectiveness of financial domain knowledge expansion and structured, logical reasoning distillation

金融AI大模型推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。